Analysis conceived and coordinated by Dr Marco Strano, with contributions by Lucia Codato on non-verbal communication and Roberto Ruggeri on forensic voice analysis and acoustic-temporal speech parameters.

Technical article by Roberto Ruggeri, forensic audio specialist. Last substantive review: 28 August 2026.

The Garlasco case is a widely known Italian judicial matter connected to Chiara Poggi and to the town of Garlasco, in northern Italy. This article does not reconstruct the case or assess personal responsibility. It uses publicly available interview material only as the technical basis for a multidisciplinary study of speech, voice and non-verbal communication.

The analysis considers pauses, hesitations, phonation time, speech density, reference-baseline differences and facial-expression indicators. None of these observations is treated as automatic evidence of deception, guilt or responsibility. The purpose is technical and educational, and the findings must be read in full respect of the presumption of innocence.

The study was conceived and coordinated by Dr Marco Strano, psychologist and criminologist. Lucia Codato, a behavioral analyst specializing in non-verbal communication, contributed the section on facial expressions and non-verbal communication. Roberto Ruggeri, forensic audio specialist and founder of audioforensicexpert.eu, contributed the acoustic-temporal speech analysis concerning phonation, pauses, hesitations, response continuity and speech density.

The study was presented in a public YouTube video published by Archivio True Crime and hosted by Gian Guido Zurli. Carlotta Barbi, author of the project Cronache di Nero Inchiostro, also took part in the live discussion. To avoid loading third-party content automatically, this page links to the video instead of embedding the YouTube player.

During the video, the general framework is explained: the criminological and methodological approach coordinated by Dr Marco Strano, the non-verbal communication analysis presented by Lucia Codato and the technical contribution by Roberto Ruggeri concerning vocal and acoustic-temporal parameters.

The remainder of this article explains the study context, the meaning of the observed indicators, the methodological limits of the analysis and the descriptive value of the measured vocal data.

Archivio True Crime — Garlasco interview analysis

The video is hosted externally on YouTube and is not loaded by this page. Selecting the button opens YouTube in a new tab.

Methodological and Legal Note

Andrea Sempio is referred to in this article exclusively as a person under investigation, in full respect of the presumption of innocence.

This article has a technical and educational purpose. It comments on audiovisual material already made public and on a study presented during a public live broadcast. The observations concern communicative, facial, vocal and acoustic-temporal parameters. They do not constitute accusations, evidence of responsibility, assessments of truthfulness, psychological or psychiatric diagnoses, or judgments of guilt.

Any judicial assessment belongs exclusively to the competent judicial authorities. The data described here must be understood as technical observations, not as evidential conclusions.

For the broader principles used to separate observation, interpretation, alternative explanations and limitations, see the Forensic Audio Methodology and Best Practice References.

The Public Interviews at the Center of the Study

During the live broadcast, it was noted that some interviews had previously been examined by the RACIS of the Carabinieri, the Italian military police scientific investigation unit, using different techniques.

The study presented in the video is not intended to replace those activities. It is a technical and educational contribution based on different observation methods concerning non-verbal communication and the paralinguistic component of speech.

The audio component described on this page is not an audio-authentication examination, a forensic speaker-comparison examination or a lie-detection procedure. Its scope is narrower: to measure and compare selected temporal properties of vocal production in defined interview segments. Questions about file integrity belong to Audio Authentication; questions about speaker attribution belong to Forensic Voice Comparison.

Roles and Contributions in the Multidisciplinary Study

Dr Marco Strano, psychologist and criminologist, conceived, coordinated and supervised the study. He structured the research and coordinated the specialist contributions.

Lucia Codato, a behavioral analyst specializing in non-verbal communication, contributed the section concerning facial-expression indicators and non-verbal communication.

Roberto Ruggeri contributed the acoustic-temporal speech section, focusing on phonation time, pauses, hesitations, response continuity, speech density and descriptive comparison with a reference segment.

The work should therefore be understood as a multidisciplinary study coordinated by Dr Marco Strano, with each specialist contributing within a distinct technical field. Simultaneous changes in different channels may justify closer descriptive examination, but they do not automatically make the methods statistically independent or increase the evidential weight of the observations.

The Study Framework Explained by Dr Marco Strano

The premise discussed by Dr Strano is that a question perceived as sensitive, critical or threatening may, in some contexts, be associated with cognitive, emotional or physiological changes. Research on deception detection, however, shows that behavioral and non-verbal cues are generally weak, variable and strongly dependent on context; no single cue can reliably establish deception or responsibility.

Within this framework, the study combines facial-action observation with acoustic-temporal speech measurements. The comparison is exploratory and descriptive. The two channels cannot be assumed a priori to provide independent evidence, and agreement between them does not automatically create stronger proof.

When several indicators change around the same passage, convergence may help identify a communicatively different moment or support further descriptive review. It does not validate the methods against each other and does not transform weak or non-specific cues into evidence of deception.

During the live broadcast, Dr Marco Strano also referred to the so-called Othello error: a truthful person who fears not being believed may display emotional signs that could be mistaken for signs of deception.

The methodological point is important. Emotional activation, fear, anger or increased cognitive effort do not reveal their cause by themselves. A person may react because of deception, but also because of pressure, fear of being misunderstood, fear of false accusation or concern that a truthful answer will not be believed.

For this reason, the indicators observed in the study cannot be interpreted automatically as signs of lying. They can identify a variation from a reference condition or a passage worthy of closer examination, but they cannot determine the cause of the change.

What Is a Reference Baseline?

A baseline is an internal comparison point. It describes how the same person speaks or behaves in a passage selected as a reference condition; it is not an absolute standard of normality and not a rigid model of how the person should speak or behave.

Each speaker shows natural within-speaker variability. Rhythm, pauses, hesitations, facial expression, posture and vocal patterns may change with topic, interview context, emotional state, fatigue, speech planning and many other factors.

Where the material permits, several reference segments selected according to documented criteria are preferable to a single passage chosen after the disputed excerpt is already known. Multiple reference samples better represent natural variability and reduce the risk of selection and context bias.

In this study, the reference comparison involved measurable temporal parameters such as phonation time, pauses, hesitations and speech density. A marked difference can be described as an acoustic-temporal deviation from the selected reference segment. It does not, by itself, explain why the deviation occurred.

The single selected reference segment used in the published example does not represent the speaker’s complete within-speaker variability. It is treated only as a descriptive comparator and not as an absolute baseline for the person.

Facial Microexpressions and Non-Verbal Communication

Lucia Codato presented the part concerning facial expressions and non-verbal communication, including rapid facial movements discussed in relation to the Facial Action Coding System (FACS), developed by Paul Ekman and Wallace V. Friesen.

FACS is a descriptive coding system that breaks facial movement into Action Units. It can document observable facial actions; it does not determine the cause of those actions and is not a stand-alone test of truthfulness.

Research specifically on microexpressions does not support simple forensic rules. Some experimental studies have reported associations under particular paradigms, while critical reviews emphasize that microexpressions and other non-verbal behaviors are not reliable, generalizable markers of deception.

Generic behaviors such as touching the face, looking in a particular direction or changing posture can be observed and described, but they cannot independently establish what a person is thinking, remembering or intending.

Acoustic-Temporal Speech Analysis

The audio component of this study concerned the description of acoustic-temporal speech parameters. It was not used to identify the speaker, authenticate the file or classify a statement as true or false.

The parameters considered included effective phonation time, pauses, hesitations, response continuity, speech density and comparison with a reference segment.

Phonation time is the duration during which vocal production is present within the selected segment. Speech density is the percentage ratio between phonation time and total segment duration:

Speech density (%) = (Phonation time / Total segment duration) × 100

When several segments are described together, a global value can be calculated by dividing the sum of phonation times by the sum of total durations. This avoids giving the same weight to excerpts of very different lengths.

Speech density is an operational descriptive measure used in this study. It is not a standardized forensic deception index and does not measure sincerity, lying, stress or responsibility.

Software Used: Praat

Praat software used for acoustic-temporal speech analysis

For the acoustic-temporal measurements, Roberto Ruggeri used Praat, phonetic-analysis software developed at the University of Amsterdam by Paul Boersma, David Weenink and collaborators.

Praat allows inspection of waveforms, spectrograms and timing information and can be used to segment phonated and non-phonated intervals. The methodological value comes from the operational definitions, segment selection and documented measurement procedure, not from the software name itself.

Comparison Between the Answer and the Reference Segment

The example discussed in the study comes from an interview broadcast by Quarto Grado on Rete 4 on 21 March 2025. The answer followed a journalistic question on a sensitive point of the case. The full question is not reproduced here; however, its interview context remains relevant to interpretation of the response pattern.

Andrea Sempio replies, in an editorial translation from Italian:

“No, that… no, no, uh, I mean, it is a… a smile of mmm… absurdity. I mean, that is it, no!”

The quotation above is an English translation of the Italian-language answer and is not a verbatim English utterance.

Waveform of the analyzed answer segment in the Garlasco interview study

The excerpt was segmented to quantify the proportion of time occupied by vocal production relative to the total duration, excluding pauses and non-phonated intervals.

Phonation and pause segmentation of the analyzed Garlasco answer

In this passage, vocal production is comparatively fragmented and includes repeated pauses and hesitations.

The data were then compared with a selected internal reference segment in which vocal production was more continuous.

Waveform of the reference segment used in the Garlasco interview analysis
Phonation and pause segmentation of the reference segment
SegmentTotal durationPhonation timeSpeech density
Answer, Quarto Grado, 21 March 202511.13 s5.09 s45.68%
Internal reference segment11.48 s9.30 s80.98%

The durations and phonation times are displayed to two decimal places. The displayed values are therefore insufficient to reproduce the percentages exactly; small differences may result from rounding of the underlying measurements.

Reported speech density comparison between the analyzed answer and the reference segment

The reported comparison shows a marked difference: speech density in the analyzed answer is 45.68%, while the selected reference segment reaches 80.98%. Descriptively, the answer therefore contains a smaller proportion of phonation and a greater proportion of pauses or non-phonated time.

This is an acoustic-temporal observation. It does not identify the cause of the difference. The same pattern may be compatible with response planning, question pressure, stress, hesitation, cognitive demand, emotion or ordinary individual discourse variability.

Research on cognitive-load approaches to deception detection mainly concerns controlled experimental paradigms or structured interviewing techniques. It does not justify inferring deception from a spontaneous reduction in speech density in a television interview.

What a Reduction in Speech Density Indicates

A lower speech-density value means that vocal production occupies a smaller proportion of the selected interval. This may result from longer pauses, hesitations, interruptions, restarts or reduced continuity of phonation.

The technical value is descriptive: instead of saying only that an answer “sounds hesitant” or “sounds fragmented”, the analysis quantifies how much of the interval contains vocal production and how much does not.

The measurement remains non-causal. Different psychological, conversational and technical conditions can produce similar temporal patterns.

What This Analysis Shows — and What It Does Not Show

Observed dataSustainable conclusionDoes not establish by itself
Speech density 45.68% in the analyzed answerA smaller proportion of phonation in that selected segmentDeception, stress or guilt
Speech density 80.98% in the reference segmentMore continuous vocal production in that selected referenceAn absolute “normal state” for the speaker
More pauses and hesitationsA descriptive difference in temporal organizationThe psychological cause of the difference
Convergence with non-verbal changesA passage that may be descriptively different across modalitiesMutual validation or proof of deception

Material and Methodological References

The measurements presented here are descriptive. This audio component is not a forensic speaker comparison and does not use the ENFSI speaker-comparison manual as a standard for inferring truthfulness or psychological state. The ENFSI manual is relevant to the broader distinction between speech-feature analysis, comparability and cautious interpretation, but its defined scope is forensic speaker comparison.

Roberto Ruggeri: Technical Contribution to the Voice Analysis

Roberto Ruggeri discussing forensic voice analysis during the public Garlasco broadcast

Roberto Ruggeri is a forensic audio specialist with more than twenty years of experience in the audio field. In this study, his contribution was limited to the acoustic-temporal speech component, with attention to phonation, pauses, hesitations, response continuity and speech density.

The contribution should not be confused with audio authentication, forensic voice comparison or psychological assessment. For professional background, see Roberto Ruggeri — forensic audio specialist. For a case-specific technical inquiry, use the contact page.


Legal and methodological note: This article analyzes publicly available communication material for technical and educational purposes. It is not a lie-detection examination, credibility assessment, psychological diagnosis or opinion on criminal responsibility. Judicial findings and procedural developments belong to the competent authorities and are outside the scope of this page.