arXiv:2409.12170cs.SDcs.AI2024-09被引 5

语音数据采集条件不统一,会导致阿尔茨海默病识别结果失真。

The Unreliability of Acoustic Systems in Alzheimer's Speech Datasets with Heterogeneous Recording Conditions

  • 用非语音部分的音频特征即可区分患者与健康人
  • 在跨语种数据集上验证了该现象的普遍性
  • 建议改用人工标注或转录文本分析,避免声学偏差

自动化语音分析是检测阿尔茨海默病早期标志物的热门方法。然而,多数阿尔茨海默病数据集的录音条件不一致,患者与对照组常在不同声学环境下评估。尽管这对基于转录或手动对齐特征的分析影响较小,但对受采集条件强烈影响的声学特征构成严重质疑。我们在广泛使用的Pitt语料库衍生的ADreSSo数据集上进行了研究,发现仅使用音频中非语音部分,基于MFCC和Wav2vec 2.0嵌入的系统仍能以高于随机水平的准确率区分患者与对照。该现象在另一组西班牙语数据集中也得到复现。因此,在这些数据集中,疾病分类可能部分由录音条件决定。本研究警示:在非标准化录音条件下,不应依赖声学系统识别患者。我们建议,对于声学异质的数据集,应仅使用转录文本或其他基于人工标注的特征进行分析,或改用严格控制声学条件收集的新数据集。

原文摘要 · Abstract (English)

Automated speech analysis is a thriving approach to detect early markers of Alzheimer's disease (AD). Yet, recording conditions in most AD datasets are heterogeneous, with patients and controls often evaluated in different acoustic settings. While this is not a problem for analyses based on speech transcription or features obtained from manual alignment, it does cast serious doubts on the validity of acoustic features, which are strongly influenced by acquisition conditions. We examined this issue in the ADreSSo dataset, derived from the widely used Pitt corpus. We show that systems based on two acoustic features, MFCCs and Wav2vec 2.0 embeddings, can discriminate AD patients from controls with above-chance performance when using only the non-speech part of the audio signals. We replicated this finding in a separate dataset of Spanish speakers. Thus, in these datasets, the class can be partly predicted by recording conditions. Our results are a warning against the use of acoustic systems for identifying patients based on non-standardized recordings. We propose that acoustically heterogeneous datasets for dementia studies should be either (a) analyzed using only transcripts or other features derived from manual annotations, or (b) replaced by datasets collected with strictly controlled acoustic conditions.

阿尔茨海默病语音分析数据偏差声学特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。