用语音合成扩充阿尔茨海默病数据,提升检测准确率
CoSTA: Cognitive-State-Conditioned TTS Data Augmentation Using ASR Transcripts for Alzheimer's Disease Detection

- 基于认知状态设计语音合成模型,生成有病理特征的语音
- 自动语音识别转写比人工转写更有效,提升数据增广效果
- 在真实数据集上实现85.83%准确率,优于已有方法
基于语音的阿尔茨海默病(AD)检测受限于病理语音数据稀缺。为此,我们提出CoSTA,一种基于文本转语音(TTS)的数据增强框架。首先,通过适配CosyVoice2和F5-TTS,构建两个认知状态条件化(CS-Cond)TTS模型,以合成具有明显阿尔茨海默病与健康对照特征的语音。此外,构建包含人工转写(MT)和36种自动语音识别(ASR)转写在内的转写语料库,研究文本来源对TTS增强的影响。同时开展增广因子分析与测试时增强实验。在ADReSS数据集上的实验表明,CS-Cond TTS显著提升合成语音实用性,且基于ASR的增强常优于人工转写驱动的增强。最终,CoSTA相较基线提升4.16%,在ADReSS测试集上达到仅音频输入下的85.83%准确率,超越现有方法。
原文摘要 · Abstract (English)
Speech-based Alzheimer's Disease (AD) detection is constrained by scarce pathological speech data. To address this, we propose CoSTA, a Text-to-Speech (TTS)-based data augmentation framework. Specifically, we first develop two Cognitive-State-Conditioned (CS-Cond) TTS models by adapting CosyVoice2 and F5-TTS to synthesize speech with distinct AD and Healthy Control characteristics. Furthermore, by constructing a transcript pool comprising Manual Transcripts (MT) and 36 Automatic Speech Recognition (ASR) transcripts, we investigate the impact of text sources on TTS-based augmentation. We also perform augmentation-factor analysis and test-time augmentation. Experiments on the ADReSS dataset show that CS-Cond TTS significantly improves synthetic speech utility, and ASR-driven augmentation frequently outperforms MT-driven augmentation. Finally, CoSTA yields a 4.16% gain over the baseline, achieving an audio-only accuracy of 85.83% on the ADReSS test set and outperforming prior methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。