Refine
Document Type
Language
- English (5)
Reviewed
Keywords
Institute
Is part of the Bibliography
- yes (5)
Going Beyond the Cookie Theft Picture Test: Detecting Cognitive Impairments Using Acoustic Features
(2022)
Automated dementia screening enables early detection and intervention, reducing costs to healthcare systems and increasing quality of life for those affected. Depression has shared symptoms with dementia, adding complexity to diagnoses. The research focus so far has been on binary classification of dementia (DEM) and healthy controls (HC) using speech from picture description tests from a single dataset. In this work, we apply established baseline systems to discriminate cognitive impairment in speech from the semantic Verbal Fluency Test and the Boston Naming Test using text, audio and emotion embeddings in a 3-class classification problem (HC vs. MCI vs. DEM). We perform cross-corpus and mixed-corpus experiments on two independently recorded German datasets to investigate generalization to larger populations and different recording conditions. In a detailed error analysis, we look at depression as a secondary diagnosis to understand what our classifiers actually learn.
For dementia screening and monitoring, standardized tests play
a key role in clinical routine since they aim at minimizing subjectivity by measuring performance on a variety of cognitive
tasks. In this paper, we report a study consisting of a semistandardized history taking followed by two standardized neuropsychological tests, namely the SKT and the CERAD-NB.
The tests include basic tasks such as naming objects, learning
word lists, but also widely used tools such as the MMSE. Most
of the tasks are performed verbally and should thus be suitable
for automated scoring based on transcripts. For the first batch
of 30 patients, we analyze the correlation between expert manual evaluations and automatic evaluations based on manual and
automatic transcriptions. For both SKT and CERAD-NB, we
observe high to perfect correlations using manual transcripts;
for certain tasks with lower correlation, the automatic scoring is
stricter than the human reference since it is limited to the audio.
Using automatic transcriptions, correlations drop as expected
and are related to recognition accuracy; however, we still observe high correlations of up to 0.98 (SKT) and 0.85 (CERADNB). We show that using word alternatives helps to mitigate
recognition errors and subsequently improves correlation with
expert scores.
Current work on speech-based dementia assessment focuses on either feature extraction to predict assessment scales, or on the automation of existing test procedures. Most research uses public data unquestioningly and rarely performs a detailed error analysis, focusing primarily on numerical performance. We perform an in-depth analysis of an automated standardized dementia assessment, the Syndrom-Kurz-Test. We find that while there is a high overall correlation with human annotators, due to certain artifacts, we observe high correlations for the severely impaired individuals, which is less true for the healthy or mildly impaired ones. Speech production decreases with cognitive decline, leading to overoptimistic correlations when test scoring
relies on word naming. Depending on the test design, fallback handling introduces further biases that favor certain groups. These pitfalls remain independent of group distributions in datasets and require differentiated analysis of target groups.
Speech pauses, alongside content and structure, offer a valuable and non-invasive biomarker for detecting dementia. This work investigates the use of pause-enriched transcripts in transformer-based language models to differentiate the cognitive states of subjects with no cognitive impairment, mild cognitive impairment, and Alzheimer’s dementia based on their
speech from a clinical assessment. We address three binary classification tasks: Onset, monitoring, and dementia exclusion. The performance is evaluated through experiments on a German Verbal Fluency Test and a Picture Description Test, comparing the model’s effectiveness across different speech production contexts. Starting from a textual baseline, we investigate the effect of incorporation of pause information and acoustic context. We show the test should be chosen depending on the task, and similarly, lexical pause information and acoustic cross-attention contribute differently.