Forensic Voice Comparison

Technical comparison between a questioned voice and reference voice samples
Forensic voice comparison is the technical examination of recordings containing a questioned voice and one or more reference voices. The purpose is to assess whether genuinely comparable features support or oppose the hypothesis that the samples originate from the same speaker.
The comparison is not based on perceived similarity alone and does not by itself provide identification with certainty. These voice comparison services depend on signal quality, usable speech, recording channel, language, speaking style and recording conditions, which determine what can be observed and how strongly the findings can be interpreted.
The work is carried out by Roberto Ruggeri, forensic audio specialist and can be handled remotely for international and cross-border matters outside Italy.
Preliminary suitability review · Preserved samples · Multidimensional examination · Documented limitations
Last substantive review: 28 August 2026
Forensic Voice Comparison and Voice Analysis: Terms and Differences

Forensic voice comparison is the term used on this page for comparison between a questioned voice and reference voice material. The ENFSI manual uses the term forensic speaker comparison for analysis of recordings containing unknown and known speakers in order to help evaluate whether the voices belong to the same or different speakers.
Voice analysis is broader. It can include acoustic, phonetic, prosodic, temporal or linguistic description of a voice even when there is no questioned sample to compare. Characterization of reference samples may support later comparison, but it is not by itself speaker attribution.
Forensic voice comparison is not necessarily the same as automatic voice biometrics and should not be confused with stress analysis, truthfulness assessment or emotion detection. Parameters such as pauses, rhythm and speech density describe vocal behavior but do not establish truth, deception or responsibility.
A separate example of acoustic-temporal voice analysis, not intended as speaker identification, is presented in the English-language Garlasco multidisciplinary case study, where phonation, pauses and hesitations are treated as descriptive data rather than automatic evidence.
When Forensic Voice Comparison May Be Requested
The service may be relevant when a recording contains a voice whose attribution is disputed and reference samples attributable to a person are available for technical comparison.
- voice messages, telephone calls or environmental conversations with disputed speaker attribution;
- anonymous or contested recordings in civil, criminal, investigative or other legal contexts;
- recordings where the first question is whether a technically meaningful comparison is feasible;
- critical review of a previous voice comparison or of the samples used for it;
- selection or acquisition of reference voice material for a later examination.
The technical question should identify which samples are being compared, which hypothesis is to be assessed, which conditions differ between the recordings and which limitations may affect interpretation.
Suitability and Comparability of Voice Samples
Before comparison, it is necessary to determine whether the material is genuinely suitable. The mere presence of two recorded voices does not automatically make a reliable comparison possible.
Amount and Distribution of Usable Speech
There is no single minimum duration that is valid for every case. What matters is the amount of clean speech actually available, phonetic variety, the number of useful occurrences and whether the same characteristics can be observed more than once.
Technical Quality and Recording Channel
Noise, reverberation, compression, clipping, telephone transmission, microphone characteristics and distance can modify the signal. Channel mismatch must be considered before comparing formants, spectral properties or voice quality.
Speaking Style, Language and Situation
Read, spontaneous, whispered, telephone or emotionally affected speech can differ substantially. Language, regional variety, phonetic content and communicative situation should be sufficiently comparable.
If the material does not meet minimum conditions for a meaningful comparison, the technically appropriate outcome may be that comparison is not practicable or that its scope must be strongly limited.
How Recorded Voices Are Compared

The sequence below is limited to speaker-comparison decisions. Cross-cutting principles for preservation, working copies, validation, quality control and reporting are maintained on the Forensic Audio Methodology page rather than repeated here in full.
1. Receipt, Identification and Working Copies
Files are identified, preserved and kept separate from the copies used for analysis. Where relevant, hashes, file format, codec, conversions and technical handling steps are documented.
2. Separate Assessment of the Samples
The questioned voice and the reference samples are first examined separately before joint comparison. This helps reduce the risk of prematurely directing observations toward expected similarities.
3. Selection of Comparable Speech
Regions with sufficient quality and useful phonetic content are selected. Unstable, heavily disturbed or non-comparable fragments are excluded or explicitly treated as limited.
4. Phonetic-Linguistic and Acoustic Analysis
Features are examined using complementary methods: critical listening, phonetic-linguistic observation, prosody, temporal organization and acoustic measurements that are appropriate to the material.
5. Evaluation of Similarities and Differences
Each convergence is considered together with incompatibilities, within-speaker variability, possible typicality of the feature and the effect of mismatched recording conditions.
6. Reporting and Conclusion
The report identifies the samples, describes the method, presents relevant findings and clarifies conditions, assumptions and limitations. The final formulation should be proportionate to the quality and actual comparability of the material.
Which Voice Characteristics May Be Examined

Phonetic-Linguistic Features
These may include articulatory habits, realization of particular speech sounds, linguistic choices and recurring patterns. Their value depends on whether they are representative, comparable and sufficiently informative in the relevant speaker population.
Prosodic and Temporal Features
Fundamental-frequency behavior, intonation, rhythm, pauses, speech rate and articulation rate can describe patterns of speech. They are sensitive to emotion, task, health, age, vocal effort and context.
Acoustic Features
Formants, voice quality and spectral characteristics may contribute where phonetic content and recording channel permit. No single measurement is treated as a unique voiceprint.
Variability and Typicality
A similarity is more informative when the feature is relatively stable within the speaker and not common in the relevant population. Where adequate population data are unavailable, the statistical scope of an observation must be expressed cautiously.
Reconstructed and Anonymized Methodological Summary Based on Professional Casework
Descriptive acoustic-phonetic characterization of several recordings attributed to the same speaker.
This reconstructed and anonymized methodological summary is based on professional casework. It was not a complete comparison between a questioned voice and a known voice, and it did not concern a deepfake. The objective was to illustrate assessment of a small reference corpus and documentation of voice characteristics that could be useful in later examinations.
To respect confidentiality, this summary does not reproduce the original report. Names, dates, locations, verbal content, exact file count, original filenames, devices, numerical values and speaker-specific graphs have been removed or generalized.
Material and Purpose
The corpus included spontaneous speech and guided reading acquired under non-uniform conditions. The purpose was to describe quality, usability, recurring features and variability without assigning identity percentages or personal attribution.
Preservation and Selection
The files received were identified and kept separate from working copies. Before measurement, noise, compression, reverberation, clipping, speaking style, tracking stability and the availability of phonetically useful regions were assessed.

Descriptive Analyses
The examination considered metadata, critical listening, spectrograms, fundamental frequency, prosody, temporal organization, recurring phonetic-linguistic patterns, stable vowel nuclei, formants and supplementary descriptors of voice quality.
Parameters were interpreted in relation to phoneme, speaking style and recording channel. Formants, jitter, shimmer, harmonic-to-noise measures and spectral profile were not treated as unique identifiers or as tests of whether a voice was natural or synthetic.

Findings and Conclusion
In technically suitable portions, recurring characteristics were observed in pitch placement, prosodic modulation, rhythmic organization, particular articulatory habits and vowel measurements that met the stated quality-control criteria.
Observed differences were interpreted in relation to different speaking tasks and technical conditions. The result was expressed as a descriptive acoustic-phonetic profile of the corpus rather than as a unique biometric voiceprint.
The profile may serve as reference material for a future comparison, but any new recording would require a new assessment of quality, amount of speech, channel, speaking style and comparability.
Reconstructed and anonymized methodological summary based on professional casework. This public account contains no transcripts, biometric values, original graphs or case-specific identifying data. Publication assumes that applicable confidentiality arrangements permit disclosure of a generalized methodological description.
How Conclusions Are Formulated
The conclusion should answer the technical question without turning comparison into an automatic verdict. Its formulation depends on the method used, availability of population information, sample quality and the level of comparability.
- similarities are not interpreted without considering how typical they may be;
- differences are assessed against natural within-speaker variation and technical mismatch;
- a software output is not treated as a stand-alone conclusion;
- percentages or numerical ratios are used only where the method and supporting data are appropriate and explicitly stated;
- where the material is insufficient, the result may be inconclusive or non-comparable.
Forensic voice comparison may therefore provide technical support for one of the competing hypotheses or identify incompatibilities, but it should not be described as absolute certainty or infallible voice recognition.
What Forensic Voice Comparison Does Not Include

This page concerns comparison between voice samples. If the technical question is different, the appropriate route may be:
Voice comparison does not establish truthfulness, emotional state or responsibility, and it does not replace legal assessment of the material.
Material Useful for a Preliminary Review
- the file containing the questioned voice in the generation closest to the source that is available;
- one or more reference voice samples with documentable provenance;
- information on device, application, recording channel and any forwarding or conversion;
- language, regional variety and manner of speech production;
- time intervals considered relevant;
- availability of the person for a controlled reference recording where appropriate;
- the technical question, case context and any relevant deadline.
Where a reference sample is acquired specifically for comparison, it is preferable to collect a substantial amount of clean speech under technical and communicative conditions that are as comparable as practicable to the questioned material. Overly rehearsed text or very different conditions can reduce comparability.
What May Be Delivered

- identification of the files and description of the material examined;
- assessment of quality, amount of usable speech and comparability;
- description of selected regions and exclusion criteria;
- relevant phonetic-linguistic, prosodic and acoustic analyses;
- tables, images and technical attachments where useful;
- discussion of similarities, differences and alternative explanations;
- a technical forensic report with conclusions and explicitly stated limitations.
The exact scope, turnaround and deliverables are agreed before the examination begins.
Frequently Asked Questions About Forensic Voice Comparison

Can forensic voice comparison identify a speaker with certainty?
No. The examination evaluates the technical support provided by genuinely comparable features and must state conditions and limitations. It does not automatically produce absolute certainty.
How much speech is needed?
There is no single number that applies to every case. Sufficient clean speech, phonetic variety and multiple usable occurrences are needed. A few seconds may be insufficient even where the recording sounds clear.
Can audio received through WhatsApp or a telephone call be used?
Potentially, but only if sufficient usable speech and technical quality remain. Duration, noise, distortion, compression, bandwidth and comparability with reference samples are assessed first. Messaging-app and telephone codecs can remove voice information and may limit the strength of the examination.
Is a specially recorded reference sample always necessary?
No. It can be useful when existing reference recordings are insufficient or unrepresentative. The recording procedure should be designed in relation to the questioned material.
Can an enhanced recording be used for comparison?
Only with caution and with untreated versions retained. Processing can alter useful characteristics and should not automatically be treated as neutral.
Do the samples need to be recorded with the same device?
Not necessarily. Differences in microphone, device, codec, acoustic environment and recording conditions may affect observable features and reduce comparability, so their impact must be considered before comparison.
Methodological References
The specific comparison framework is grounded in the ENFSI Best Practice Manual for the Methodology of Forensic Speaker Comparison, FSA-BPM-003, Issue 001. The manual describes a combined phonetic-linguistic auditory and acoustic approach and explicitly addresses mismatched conditions between questioned and reference recordings.
For general handling, preservation, working copies and forensic-audio documentation, see the SWGDE Best Practices for Forensic Audio.
Request a Preliminary Review of the Voice Samples

Send the version closest to the source of the questioned voice, the available reference samples and a brief description of the context. The preliminary review is used to assess suitability, comparability and the main technical limitations before a full examination is considered.