Refine
Document Type
- conference proceeding (article) (7)
- Article (1)
- Part of a Book (1)
- Report (1)
- Working Paper (1)
Reviewed
- Begutachtet/Reviewed (10)
Keywords
- neuropsychological tests (3)
- dementia screening (2)
- pathological speech (2)
- Autonomes Fahren (1)
- ChatBot (1)
- Gesichtserkennung (1)
- Human-Machine Improvisation, Co-creativity, Player Piano (1)
- Künstliche Intelligenz (1)
- Personal Narratives, Emotion Annotation, Segment Level Annotation (1)
- dementia assessment (1)
Institute
Going Beyond the Cookie Theft Picture Test: Detecting Cognitive Impairments Using Acoustic Features
(2022)
Automated dementia screening enables early detection and intervention, reducing costs to healthcare systems and increasing quality of life for those affected. Depression has shared symptoms with dementia, adding complexity to diagnoses. The research focus so far has been on binary classification of dementia (DEM) and healthy controls (HC) using speech from picture description tests from a single dataset. In this work, we apply established baseline systems to discriminate cognitive impairment in speech from the semantic Verbal Fluency Test and the Boston Naming Test using text, audio and emotion embeddings in a 3-class classification problem (HC vs. MCI vs. DEM). We perform cross-corpus and mixed-corpus experiments on two independently recorded German datasets to investigate generalization to larger populations and different recording conditions. In a detailed error analysis, we look at depression as a secondary diagnosis to understand what our classifiers actually learn.
In recent years, speech processing for medical applications got significant traction. While pioneering work in the 1990ies focused on processing sustained vowels or isolated utterances, work in the 2000s already showed, that speech recognition systems, prosodic analysis and natural language processing be used to assess a large variety of speech pathologies.Here, we give an overview of how to classify selected speech pathologies including stuttering, language development, speech intelligibility after surgery, dementia and Alzheimers, depression and state-of-mind. While each of those poses a rather well-defined problem in a lab setting, we discuss the issues when integrating such methods in a clinical workflow such as diagnosis or monitoring. Starting from the question if such detectors can be used for general screening or rather as a specialist's tool, we explore the legal and privacy-related implications: patient-doctor conversations, working with children or demented seniors, bias towards examiner or patient, on-device vs. cloud processing.We conclude with a set of open questions that should be addressed to help bringing all this research from the lab to routine clinical use.
The detection of pathologies from speech features is usually defined as a binary classification task with one class representing a specific pathology and the other class representing healthy speech. In this work, we train neural networks, large margin classifiers, and tree boosting machines to distinguish between four pathologies: Parkinson's disease, laryngeal cancer, cleft lip and palate, and oral squamous cell carcinoma. We show that latent representations extracted at different layers of a pre-trained wav2vec 2.0 system can be effectively used to classify these types of pathological voices. We evaluate the robustness of our classifiers by adding room impulse responses to the test data and by applying them to unseen speech corpora. Our approach achieves unweighted average F1-Scores between 74.1% and 97.0%, depending on the model and the noise conditions used. The systems generalize and perform well on unseen data of healthy speakers sampled from a variety of different sources.
Künstliche Intelligenz
(2018)
Dieses Arbeitspapier wurde von Studierenden der Fakultät Betriebswirtschaft im Rahmen eines Seminars zur Künstlichen Intelligenz erarbeitet. Darin entwickeln die Studierenden einen historischen Überblick über das Feld der Künstlichen Intelligenz, wichtige Fachgebiete und stellen dann aktuelle Anwendungen wie Gesichts- oder Spracherkennung und autonomes Fahren vor. Außerdem dokumentieren die Studierenden ihre eigenen Erfahrungen in der Erstellung eines ChatBots, der auf Verfahren der künstlichen Intelligenz mithilfe von IBM Watson erstellt wurde.
For dementia screening and monitoring, standardized tests play
a key role in clinical routine since they aim at minimizing subjectivity by measuring performance on a variety of cognitive
tasks. In this paper, we report a study consisting of a semistandardized history taking followed by two standardized neuropsychological tests, namely the SKT and the CERAD-NB.
The tests include basic tasks such as naming objects, learning
word lists, but also widely used tools such as the MMSE. Most
of the tasks are performed verbally and should thus be suitable
for automated scoring based on transcripts. For the first batch
of 30 patients, we analyze the correlation between expert manual evaluations and automatic evaluations based on manual and
automatic transcriptions. For both SKT and CERAD-NB, we
observe high to perfect correlations using manual transcripts;
for certain tasks with lower correlation, the automatic scoring is
stricter than the human reference since it is limited to the audio.
Using automatic transcriptions, correlations drop as expected
and are related to recognition accuracy; however, we still observe high correlations of up to 0.98 (SKT) and 0.85 (CERADNB). We show that using word alternatives helps to mitigate
recognition errors and subsequently improves correlation with
expert scores.
In recent years, machine learning, and in particular generative adversarial neural networks (GANs) and attention-based neural networks (transformers), have been successfully used to compose and generate music, both melodies and polyphonic pieces. Current research focuses foremost on style replication (e.g., generating a Bach-style chorale) or
style transfer (e.g., classical to jazz) based on large amounts of recorded or transcribed music, which in turn also allows for fairly straight-forward “performance” evaluation. However, most of these models are not suitable for human-machine co-creation through live interaction, neither is clear, how such models and resulting creations would be evaluated.
This article presents a thorough review of music representation, feature analysis, heuristic algorithms, statistical and parametric modelling, and human and automatic evaluation measures, along with a discussion of which approaches and models seem most suitable for live interaction.
Spirio Sessions
(2021)
This paper presents an ongoing interdisciplinary research project that deals with free improvisation and human-machine interaction, involving a digital player piano and other musical instruments. Various technical concepts are developed by student participants in the project and continuously evaluated in artistic performances. Our goal is to explore methods for co-creative collaborations with artificial intelligences embodied in the player piano, enabling it to act as an equal improvisation partner for human musicians.
Personal Narrative (PN) is the recollection of individuals’ life experiences, events, and thoughts along with the associated emotions in the form of a story. Compared to other genres such as social media texts or microblogs, where people write about ex-perienced events or products, the spoken PNs are complex to analyze and understand. They are usually long and unstructured, involving multiple and related events, characters as well as thoughts and emotions associated with events, objects, and persons. In spoken PNs, emotions are conveyed by changing the speech signal characteristics as well as the lexical content of the narrative. In this work, we annotate a corpus of spoken personal narratives, with the emotion valence using discrete values. The PNs are segmented into speech segments, and the annotators annotate them in the discourse context, with values on a 5 point bipolar scale ranging from -2 to +2 (0 for neutral). In this way, we capture the unfolding of the PNs events and changes in the emotional state of the narrator. We perform an in-depth analysis of the inter-annotator agreement, the relation between the label distribution w.r.t. the stimulus (positive/negative) used for the elicitation of the narrative, and compare the segment-level annotations to a baseline continuous annotation. We find that the neutral score plays an important role in the agreement. We observe that it is easy to differentiate the positive from the negative valence while the confusion with the neutral label is high.
Current work on speech-based dementia assessment focuses on either feature extraction to predict assessment scales, or on the automation of existing test procedures. Most research uses public data unquestioningly and rarely performs a detailed error analysis, focusing primarily on numerical performance. We perform an in-depth analysis of an automated standardized dementia assessment, the Syndrom-Kurz-Test. We find that while there is a high overall correlation with human annotators, due to certain artifacts, we observe high correlations for the severely impaired individuals, which is less true for the healthy or mildly impaired ones. Speech production decreases with cognitive decline, leading to overoptimistic correlations when test scoring
relies on word naming. Depending on the test design, fallback handling introduces further biases that favor certain groups. These pitfalls remain independent of group distributions in datasets and require differentiated analysis of target groups.