Science

Every technical claim, with the strength of evidence it actually has

We organize the bibliography by technology pillar. Where the science is solid, we show the study and the exact number. Where it's a proprietary framework, we say so without hedging.

Evidence explorer

Strong evidence

Depression and psychomotor retardation

Less variable voice in F0, more pauses, and slower speech in depressive presentations — replicated across multiple independent studies.

Zhao, Q. et al. (2022). Vocal Acoustic Features as Potential Biomarkers for Identifying/Diagnosing Depression: A Cross-Sectional Study. Frontiers in Psychiatry. 71 depressed patients, 62 controls. MFCC7 predictive of PHQ-9 (β=0.90, p=0.01); MFCC9 correlated with HAMD somatic anxiety (r=−0.34, β=−0.45, p=0.049); 89.66% discriminant accuracy. DOI 10.3389/fpsyt.2022.815678 — full text checked directly at the source.
Taguchi, T. et al. (2018). Major depressive disorder discrimination using vocal acoustic features. Journal of Affective Disorders. MFCC2 discriminates MDD with 77.8% sensitivity / 86.1% specificity. DOI 10.1016/j.jad.2017.08.038
Quatieri, T. & Malyska, N. (2012). Vocal-Source Biomarkers for Depression: A Link to Psychomotor Activity. Interspeech. DOI 10.21437/interspeech.2012-311

Rigor caveat: the same patterns (reduced F0, flattened prosody) also appear in Parkinson's disease, frontotemporal degeneration, ALS, and schizophrenia. It is not a marker specific to depression alone — it is a transdiagnostic signal of neuromotor slowing, and we treat it as such.

Strong evidence

MFCC and temporal derivatives (Δ, ΔΔ)

Static cepstral coefficients and their derivatives capture temporal dynamics relevant to affective state and articulatory coordination.

Williamson, J. et al. (2013). Vocal biomarkers of depression based on motor incoordination. Proc. 3rd ACM AVEC Workshop. DOI 10.1145/2512530.2512531
Davis, S. & Mermelstein, P. (1980). Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences. IEEE Trans. ASSP. — conceptual foundation of MFCCs, cited directly in FROID's engine.

Caveat: deltas do not always outperform the static MFCC — it depends on the task and the classifier.

Emerging evidence

Prosody in mania and bipolarity

Dynamic F0 and speech rate change with mood state, but with high intra- and inter-individual variability.

Guidi, A. et al. (2015). Automatic analysis of speech F0 contour for the characterization of mood changes in bipolar patients. Biomedical Signal Processing and Control. DOI 10.1016/j.bspc.2014.10.011
Anmella, G. et al. (2024). Automated Speech Analysis in Bipolar Disorder: The CALIBER Study Protocol and Preliminary Results. Journal of Clinical Medicine. Preliminary/protocol results. DOI 10.3390/jcm13174997

Scientific honesty: Pal et al. (2025), Annals of Indian Psychiatry, found no significant difference in F0, intensity, jitter, or shimmer between mania, depression, and euthymia (DOI 10.4103/aip.aip_221_24). We include this contrary result on purpose.

Strong infrastructure Dissimulation: contested

Facial dynamics (FACS + Markov models)

Onset/apex/offset modeling of Action Units with HMM/HSMM is a mature methodology in computer vision.

Valstar, M. & Pantic, M. (2012). Fully Automatic Recognition of the Temporal Phases of Facial Actions. IEEE TSMC-B. DOI 10.1109/tsmcb.2011.2163710
Hamm, J. et al. (2011). Automated Facial Action Coding System for dynamic analysis of facial expressions in neuropsychiatric disorders. Journal of Neuroscience Methods, 200(2), 237–256.
Kawulok, M. et al. (2021). Dynamics of facial actions for assessing smile genuineness. PLoS ONE. DOI 10.1371/journal.pone.0244647

Important caveat: the classic "Duchenne smile" hypothesis (AU6+AU12 = genuine smile) is contested by recent research, which found that AU6 correlates with smile intensity, not emotional authenticity. FROID uses this configuration as one factor among several, not as a standalone dissimulation detector.

Weak evidence

Vocal micro-tremor (5–12 Hz) and autonomic activation

There is real literature on physiological vocal microtremor and on F0/jitter/shimmer reacting to stress in general.

Schoentgen, J. (2002). Modulation frequency and modulation level owing to vocal microtremor. Journal of the Acoustical Society of America. DOI 10.1121/1.1492820 — characterizes microtremor in normal/dysphonic voices; does not establish the specific autonomic/psychiatric link that FROID monitors.
De Lacerda Veiga, D. et al. (2025). The Fundamental Frequency of Voice as a Potential Stress Biomarker: A Systematic Review and Meta-Analysis. Stress and Health. High heterogeneity, insufficient evidence for standalone clinical use. DOI 10.1002/smi.70112

This is the system's most experimental pillar — we treat the subharmonic_energy_5_12hz indices as an exploratory signal, not a validated biomarker.

No published scientific basis Proprietary framework

The 12 Perception Zones

The physics behind it (power spectral density via FFT) is real. The mapping of specific Hz ranges to emotional dichotomies has no peer-reviewed source — it follows the solfeggio/chakra frequency pattern, and its real origin is commercial voice biofeedback technology.

We communicate this explicitly as a FROID product heuristic — never as validated neuroscience.

Verification methodology

How we arrived at this bibliography

Every citation on this page was independently checked — DOI resolved, and in the case of the anchor study (Zhao et al., 2022), the full text was read at the original source to confirm that the numbers match exactly what FROID cites.