Illustration of human hearing, frequency sensitivity and sound perception

Technical guide by Roberto Ruggeri, forensic audio specialist. Last substantive review: 27 August 2026.

Human hearing is the process by which acoustic pressure changes are converted into neural signals and perceived as sound. The auditory system does not reproduce the physical waveform in a simple or linear way: frequency, level, duration, masking, binaural cues and the listener’s own hearing all influence what becomes audible and how it is perceived.

Understanding human hearing and sound perception therefore requires separating the physical properties of a sound from their perceptual correlates. Frequency is not identical to pitch, sound-pressure level is not identical to loudness, and the physical presence of acoustic energy does not guarantee that a listener can distinguish it in noise.

This guide explains how human hearing works, from the outer ear to the brain, and then examines frequency sensitivity, loudness, masking, temporal resolution, spatial hearing and speech perception in noise. The final section explains why these distinctions matter when recorded sound is interpreted in a forensic context.

How Human Hearing Works: From Sound Wave to Perception

The hearing process begins when pressure waves in air reach the outer ear. The National Institute on Deafness and Other Communication Disorders (NIDCD) describes a sequence in which sound travels through the ear canal, moves the eardrum and middle-ear ossicles, produces fluid motion in the cochlea and is converted by sensory hair cells into electrical activity carried by the auditory nerve.

StageMain function
Outer earCollects sound and directs pressure variations through the ear canal toward the eardrum.
Middle earTransfers vibration through the malleus, incus and stapes toward the inner ear.
CochleaMechanical vibration produces a traveling wave along the basilar membrane; different cochlear regions respond preferentially to different frequency regions.
Hair cells and auditory nerveSensory hair-cell motion is converted into neural activity that is transmitted toward the brain.
Central auditory systemNeural information is progressively analyzed and integrated into percepts such as pitch, loudness, location, timbre and meaningful sound.

The cochlea is tonotopically organized: higher-frequency stimulation is represented mainly toward the basal end, while lower-frequency stimulation produces its strongest mechanical response farther toward the apex. This frequency organization continues through several stages of the auditory pathway.

Physical Sound and Sound Perception Are Different

Audio analysis describes measurable signal properties. Hearing adds a perceptual layer. The relationship between the two is systematic, but not one-to-one.

Physical / acoustic propertyRelated perceptual propertyImportant limitation
Frequency and spectral structurePitch and timbreComplex sounds are not perceived from one frequency value alone.
Sound pressure / signal levelLoudnessEqual physical levels at different frequencies do not necessarily sound equally loud.
Duration, onset, envelope and temporal gapsTiming, rhythm and event separationVery brief or closely spaced events can be difficult to separate perceptually.
Differences between signals at the two earsSpatial locationNoise, reverberation, playback geometry and hearing status can degrade spatial cues.

This distinction is central to psychoacoustics: the acoustic signal constrains perception, but the listener’s auditory system determines how that signal is represented and experienced.

Human Hearing Range and Frequency Sensitivity

The often-cited range of approximately 20 Hz to 20 kHz is a useful broad description of young, normal human hearing under favorable conditions. It is not a fixed range that applies identically to every person or listening situation. Thresholds vary with frequency, age, exposure history, hearing status, presentation method and sound level.

Hearing sensitivity is also far from uniform across that range. The threshold for detecting a low-frequency tone can be much higher than the threshold for a tone in a more sensitive mid-frequency region. This is one reason a simple statement such as “the recording contains energy at this frequency” does not by itself establish how audible that component was to a particular listener.

Frequency and Pitch Are Related but Not Identical

Frequency is a physical property measured in hertz. Pitch is a perceptual attribute describing how high or low a sound seems. For a simple pure tone, frequency and pitch are closely related. Real-world sounds, however, usually contain multiple spectral components and time-varying structure.

For complex sounds such as speech, musical notes and environmental noise, perceived pitch can depend on harmonic relationships and temporal information rather than on a single spectral peak. A review by Moore (2008) describes how auditory frequency selectivity and temporal processing contribute to the analysis of speech sounds.

Why Loudness Is Not the Same as Decibels

Loudness is perceptual; sound-pressure level is physical. Increasing sound pressure generally increases perceived loudness, but the relationship depends on frequency and listening conditions.

ISO 226:2023 specifies combinations of frequency and sound-pressure level that are perceived as equally loud under defined conditions: binaural listening to frontal pure tones by otologically normal listeners aged 18–25 in a specified sound field.

The resulting equal-loudness contours demonstrate an important principle: two tones presented at the same physical sound-pressure level can be perceived at different loudnesses. The standard should not be treated as a universal prediction for every individual, headphone, room, complex signal or age group; its conditions are part of the result.

Auditory Masking: Why a Present Sound Can Be Hard to Hear

Auditory masking occurs when the presence of one sound raises the detection threshold or reduces the perceptual accessibility of another. A target sound can therefore be physically present in a signal yet difficult to detect or distinguish when competing sound is sufficiently strong or acoustically similar.

Masking is not one single mechanism. Two useful distinctions in speech and complex listening are:

TypeWhat interferesExample
Energetic maskingAcoustic energy from the masker overlaps or competes with target information within the auditory system.Steady noise obscuring weak speech components in the same frequency region.
Informational maskingCompeting perceptual or linguistic information makes it harder to segregate or attend to the target even beyond peripheral energetic overlap.Trying to follow one talker while another intelligible talker is speaking nearby.

Reviews of speech perception in noise show why the relationship between target, masker and listener is more informative than a simple statement that “noise is present”.

Temporal Resolution and Short Acoustic Events

Hearing also depends on time. The auditory system follows changes in onset, offset, amplitude envelope, modulation and short gaps, but temporal resolution is finite. Two events that are physically separate can become difficult to distinguish if they are sufficiently brief, closely spaced, masked or reverberant.

This matters for transient sounds and rapid speech cues. A sharp recorded event may be easy to locate in a waveform or spectrogram while its perceptual character depends on duration, surrounding sound and playback conditions. Conversely, a brief gap or low-level component may be measurable without being reliably salient to a listener.

Binaural Hearing and Sound Localization

With two ears, the auditory system can compare differences between the signals arriving at each side of the head. Important spatial cues include interaural time differences (ITD), interaural level differences (ILD) and frequency-dependent spectral effects produced by the head and outer ears.

A 2024 review of auditory localization describes how these binaural and monaural cues work together with factors such as reverberation and motion. Spatial hearing is therefore not equivalent to comparing volume alone between the left and right channels of a recording.

The recording and playback chain can also change spatial information. A mono file, single-channel microphone, distant source, reverberant room, headphone playback or speaker placement can each limit how closely the reproduced listening situation corresponds to the original acoustic scene.

Speech Perception in Noise: Audibility Is Not Intelligibility

A listener may detect that speech is present without understanding the words. Conversely, some speech cues may remain usable even when much of the signal is masked. Speech perception in noise depends on signal-to-noise relationship, spectral and temporal overlap, competing talkers, language-related information and the listener’s auditory function.

This distinction prevents a common error: “audible” and “intelligible” are not synonyms. If the practical question is whether processing can improve an existing recording, that belongs to the separate topic of Forensic Audio Restoration, where enhancement and its limitations are addressed directly.

Age, Noise Exposure and Individual Variability

Human hearing varies substantially between listeners. NIDCD guidance on age-related hearing loss notes that hearing can change gradually with age and that changes in the inner ear, auditory pathways, long-term noise exposure and other factors can contribute. Difficulty following speech in noise can be one practical consequence.

Noise exposure is another source of variability. NIOSH identifies occupational exposure to hazardous noise as a preventable cause of permanent hearing loss. Exposure effects depend on level, duration and history rather than on job title alone.

For this reason, age or occupation cannot by themselves establish what a specific person could hear. An individual conclusion requires information about the listener and the actual listening conditions rather than population averages alone.

Why Hearing Science Matters in Forensic Audio

Forensic audio often requires a clear distinction between what exists in the recorded signal and what a listener could reasonably perceive under specified conditions. Hearing science helps define that distinction, but it does not replace case-specific evidence.

  • A measurable frequency component is not automatically equally audible at every frequency.
  • A physically present voice component may be masked by stronger sound.
  • Detection of speech does not establish reliable word intelligibility.
  • A perceived source direction depends on binaural and spectral cues, not on channel level alone.
  • Population hearing curves cannot determine the hearing ability of a specific individual.
  • A screenshot or spectrogram cannot by itself establish what a person heard in the original acoustic environment.

When the issue is visual time-frequency interpretation rather than perception, see How to Read a Spectrogram in Forensic Audio. When the issue is whether ambiguous sound becomes a specific word or voice because of suggestion or expectation, see Auditory Pareidolia in Audio Recordings. General principles for documenting observations, interpretations and limitations are described in the Forensic Audio Methodology and Best Practice References.

Scientific and Technical References


Scope note: This article explains general auditory physiology and psychoacoustics. It does not diagnose hearing loss, determine what a specific person heard in a particular event or provide a legal conclusion. Individual audibility or intelligibility questions require the actual recording, playback conditions and relevant listener information.