Science
We organize the bibliography by technology pillar. Where the science is solid, we show the study and the exact number. Where it's a proprietary framework, we say so without hedging.
Evidence explorer
Less variable voice in F0, more pauses, and slower speech in depressive presentations — replicated across multiple independent studies.
Rigor caveat: the same patterns (reduced F0, flattened prosody) also appear in Parkinson's disease, frontotemporal degeneration, ALS, and schizophrenia. It is not a marker specific to depression alone — it is a transdiagnostic signal of neuromotor slowing, and we treat it as such.
Static cepstral coefficients and their derivatives capture temporal dynamics relevant to affective state and articulatory coordination.
Caveat: deltas do not always outperform the static MFCC — it depends on the task and the classifier.
Dynamic F0 and speech rate change with mood state, but with high intra- and inter-individual variability.
Scientific honesty: Pal et al. (2025), Annals of Indian Psychiatry, found no significant difference in F0, intensity, jitter, or shimmer between mania, depression, and euthymia (DOI 10.4103/aip.aip_221_24). We include this contrary result on purpose.
Onset/apex/offset modeling of Action Units with HMM/HSMM is a mature methodology in computer vision.
Important caveat: the classic "Duchenne smile" hypothesis (AU6+AU12 = genuine smile) is contested by recent research, which found that AU6 correlates with smile intensity, not emotional authenticity. FROID uses this configuration as one factor among several, not as a standalone dissimulation detector.
There is real literature on physiological vocal microtremor and on F0/jitter/shimmer reacting to stress in general.
This is the system's most experimental pillar — we treat the subharmonic_energy_5_12hz indices as an exploratory signal, not a validated biomarker.
The physics behind it (power spectral density via FFT) is real. The mapping of specific Hz ranges to emotional dichotomies has no peer-reviewed source — it follows the solfeggio/chakra frequency pattern, and its real origin is commercial voice biofeedback technology.
We communicate this explicitly as a FROID product heuristic — never as validated neuroscience.
Verification methodology
Every citation on this page was independently checked — DOI resolved, and in the case of the anchor study (Zhao et al., 2022), the full text was read at the original source to confirm that the numbers match exactly what FROID cites.