Most cognitive assessment today happens in snapshots: a brief in-clinic test such as the MMSE or MoCA, a clinical interview, and — if something looks off — a longer neuropsychological battery, often a year or more later. These tools are valuable and validated. But they share a structural limitation: they are episodic, and cognition doesn't decline on an annual schedule.
The scale of the problem makes that gap matter. Global dementia cases are projected to rise from 57.4 million in 2019 to 152.8 million by 2050. Detecting meaningful change earlier — and monitoring it over time — is one of the field's central challenges.
What "longitudinal" actually buys you
A single score tells you where someone sits relative to a population average. It's far less good at catching the first, subtle deviation from their own normal — which may be the earliest signal of all. Tracking change within a person, over time, is a different measurement problem, and the published evidence suggests language and speech carry signal that unfolds across exactly those timescales.
The Nun Study famously found that linguistic ability measured in early adulthood was associated with cognitive function and Alzheimer's decades later. More recently, researchers using the Framingham cohort showed that language features could help predict future onset of Alzheimer's years ahead of diagnosis. And in autopsy-confirmed Alzheimer's, connected-speech features tracked disease progression over time. Different studies, one throughline: speech and language change are measurable, and they move with the underlying trajectory.
Why speech, and why now
Speech is uniquely suited to frequent measurement. It's natural, low-friction, and repeatable — you can capture it briefly and often, remotely, without a clinic visit. That's the premise behind the growing interest in digital biomarkers from everyday devices. Groups have shown that automated remote speech assessment can detect early signs of cognitive impairment, and that a smartphone-based speech test can screen for early Alzheimer's against an amyloid-confirmed cohort.
Crucially, the methodology is mature enough to be rigorous. Standardized acoustic feature sets like eGeMAPS and open shared-task benchmarks like ADReSS mean approaches can be compared rather than taken on faith.
The honest limits
None of this means the problem is solved. Most published results are cross-sectional group classification — distinguishing people who already have a diagnosis from those who don't, at a single point in time — rather than within-person tracking. Performance varies with task, language, and recording conditions, and larger prospective validation is repeatedly called for. A 2025 meta-analysis put pooled accuracy for MCI around 80%, but with wide confidence intervals. Continuous, longitudinal cognitive monitoring is a promising hypothesis — not a finished science.
Key takeaways
- Standard cognitive tests are episodic; decline is continuous. That mismatch can hide early change.
- Published work shows speech and language change measurably — and, in several studies, longitudinally and even predictively.
- Speech is low-friction and repeatable, making frequent remote measurement plausible.
- But most evidence is cross-sectional group classification, not within-person tracking — so this remains a hypothesis under evaluation.
How Recalla approaches it
Recalla's bet follows directly from the evidence and its limits: frequent, low-friction voice check-ins that build a personal baseline, so the signal we care about is a person's change from their own normal — as a complement to, not a replacement for, clinical assessment. Every output is explainable and non-diagnostic, and our public demo runs on synthetic data. Whether this is clinically useful in the real world is exactly what we're setting out to test.
For the full picture of what published science establishes — and what Recalla has and hasn't validated — see the Evidence Hub.
References
- GBD 2019 Dementia Forecasting Collaborators (2022). Estimation of the global prevalence of dementia in 2019 and forecasted prevalence in 2050. The Lancet Public Health, 7(2):e105–e125. doi:10.1016/S2468-2667(21)00249-8
- Snowdon, D.A. et al. (1996). Linguistic Ability in Early Life and Cognitive Function and Alzheimer's Disease in Late Life. JAMA, 275(7):528–532. doi:10.1001/jama.1996.03530310034029
- Eyigoz, E. et al. (2020). Linguistic markers predict onset of Alzheimer's disease. eClinicalMedicine, 28:100583. doi:10.1016/j.eclinm.2020.100583
- Ahmed, S. et al. (2013). Connected speech as a marker of disease progression in autopsy-proven Alzheimer's disease. Brain, 136(12):3727–3737. doi:10.1093/brain/awt269
- Kourtis, L.C. et al. (2019). Digital biomarkers for Alzheimer's disease: the mobile/wearable devices opportunity. npj Digital Medicine, 2:9. doi:10.1038/s41746-019-0084-2
- Robin, J. et al. (2021). Using Digital Speech Assessments to Detect Early Signs of Cognitive Impairment. Frontiers in Digital Health, 3:749758. doi:10.3389/fdgth.2021.749758
- Fristed, E. et al. (2022). A remote speech-based AI system to screen for early Alzheimer's disease via smartphones. Alz. & Dementia: DADM, 14:e12366. doi:10.1002/dad2.12366
- Eyben, F. et al. (2016). The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affective Computing. IEEE Trans. Affective Computing, 7(2):190–202. doi:10.1109/TAFFC.2015.2457417
- Luz, S. et al. (2020). Alzheimer's Dementia Recognition Through Spontaneous Speech: The ADReSS Challenge. Interspeech 2020, 2172–2176. doi:10.21437/Interspeech.2020-2571
- de la Fuente Garcia, S., Ritchie, C.W. & Luz, S. (2020). AI, Speech, and Language Processing Approaches to Monitoring Alzheimer's Disease: A Systematic Review. J. Alzheimer's Disease, 78(4):1547–1574. doi:10.3233/JAD-200888
- Jafari, Z., Andrew, M.K. & Rockwood, K.J. (2025). Diagnostic utility of speech-based biomarkers in mild cognitive impairment: a systematic review and meta-analysis. Age and Ageing, 54(10):afaf316. doi:10.1093/ageing/afaf316