arXiv:2508.20513cs.SDcs.MM2025-08被引 2

用语音合成增强数据,结合专家模型选特征,提升阿尔茨海默病早期筛查准确率

MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening

  • 通过文本转语音合成扩充数据,解决样本不足问题
  • 引入混合专家机制动态选择关键声学与文本特征,分类准确率达85.71%
  • 适合数据稀缺场景下的临床辅助诊断系统开发

通过语音进行阿尔茨海默病(AD)的早期筛查是一种有前景的无创方法。然而,数据有限及缺乏细粒度、自适应特征选择常限制性能。为此,我们提出MoTAS框架,旨在提升AD筛查效率。该框架利用文本转语音(TTS)增强数据量,并采用混合专家(MoE)机制优化多模态特征选择,共同提升模型泛化能力。流程始于自动语音识别(ASR)获取精准转录;随后使用TTS合成语音以丰富数据集;提取声学与文本嵌入后,通过MoE机制动态筛选最具信息量的特征,优化特征融合以提升分类效果。在ADReSSo数据集上的评估显示,MoTAS达到85.71%的领先准确率,优于现有基线。消融实验进一步验证了TTS增强与MoE机制对性能提升的独立贡献。研究结果表明,MoTAS在真实世界的数据受限场景下具有重要应用价值。

原文摘要 · Abstract (English)

Early screening for Alzheimer's Disease (AD) through speech presents a promising non-invasive approach. However, challenges such as limited data and the lack of fine-grained, adaptive feature selection often hinder performance. To address these issues, we propose MoTAS, a robust framework designed to enhance AD screening efficiency. MoTAS leverages Text-to-Speech (TTS) augmentation to increase data volume and employs a Mixture of Experts (MoE) mechanism to improve multimodal feature selection, jointly enhancing model generalization. The process begins with automatic speech recognition (ASR) to obtain accurate transcriptions. TTS is then used to synthesize speech that enriches the dataset. After extracting acoustic and text embeddings, the MoE mechanism dynamically selects the most informative features, optimizing feature fusion for improved classification. Evaluated on the ADReSSo dataset, MoTAS achieves a leading accuracy of 85.71\%, outperforming existing baselines. Ablation studies further validate the individual contributions of TTS augmentation and MoE in boosting classification performance. These findings highlight the practical value of MoTAS in real-world AD screening scenarios, particularly in data-limited settings.

阿尔茨海默病语音分析多模态学习TTS增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。