Acoustic noise alters speech-based Alzheimer's model predictions
Across three self-supervised learning models tested on ADReSSo speech audio, controlled acoustic interventions systematically shifted diagnostic predictions for Alzheimer's disease.
In a preprint analyzing human speech recordings from the ADReSSo dataset, researchers investigated whether acoustic factors influence speech-based Alzheimer's disease assessments. They evaluated three pretrained self-supervised learning backbones using layer-wise linear decoding, input- and representation-space interventions, and geometric alignment analysis. The authors applied controlled noise and reverberation interventions to participant-speech-only, non-speech, and full-recording audio. Controlled acoustic interventions altered diagnostic predictions across all three backbones. Background noise produced the strongest intervention effects, despite showing no significant diagnostic-group difference in the original dataset. These shifts aligned systematically with the classifiers' decision directions, replicated on a held-out test set, and reversed when representation-space interventions were reversed.
Why it matters
Acoustic biomarkers are emerging tools for detecting age-related cognitive impairment. These findings indicate that models can rely on recording artifacts rather than genuine disease signs, risking misclassification in clinical settings.
Caveats
The findings are based on a computational analysis limited to the ADReSSo dataset and three specific model architectures. The study is a preprint and has not yet undergone peer review.
The paper
Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?
Kopar S, Koudounas A, Rane RP et al.
arXiv · 1 Oct 2026 · Preprint, not yet peer-reviewed

