用分层二分类与语音融合提升痴呆早期检测准确率
Leveraging Cascaded Binary Classification and Multimodal Fusion for Dementia Detection through Spontaneous Speech
- 分层二分类框架,结合停顿编码捕捉语言不流畅性
- 多模态融合模型在认知评分预测中超越基线
- 适合临床辅助诊断与自然语言处理研究者参考
本文提交至PROCESS Challenge 2025,聚焦自发语音分析以实现痴呆早期检测。针对三分类任务(健康对照、轻度认知障碍、痴呆),提出一种级联二分类框架,微调预训练语言模型并引入停顿编码,更有效捕捉语言不流畅性,简化多分类流程并缓解类别不平衡问题。对于简易精神状态检查(MMSE)得分回归任务,构建增强型多模态融合系统,整合多样声学与语言特征;对各特征集分别训练回归模型,并通过得分平均实现集成学习。在测试集上,两项任务均优于主办方提供的基线模型,验证了方法的鲁棒性与有效性。
原文摘要 · Abstract (English)
This paper presents our submission to the PROCESS Challenge 2025, focusing on spontaneous speech analysis for early dementia detection. For the three-class classification task (Healthy Control, Mild Cognitive Impairment, and Dementia), we propose a cascaded binary classification framework that fine-tunes pre-trained language models and incorporates pause encoding to better capture disfluencies. This design streamlines multi-class classification and addresses class imbalance by restructuring the decision process. For the Mini-Mental State Examination score regression task, we develop an enhanced multimodal fusion system that combines diverse acoustic and linguistic features. Separate regression models are trained on individual feature sets, with ensemble learning applied through score averaging. Experimental results on the test set outperform the baselines provided by the organizers in both tasks, demonstrating the robustness and effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。