用语音分析检测认知衰退,融合多模态特征提升诊断准确率
Tackling Cognitive Impairment Detection from Speech: A submission to the PROCESS Challenge
- 结合声学、文本与大模型语言特征,构建多维度语音分析体系
- 在三个临床任务中综合表现最佳,模型组合实现互补性提升
- 适合医疗AI、认知健康研究者参考,助力早期认知障碍筛查
本工作描述了我们团队参与2024年PROCESS挑战赛的提交方案,目标是通过自发性语音评估认知衰退,涵盖三项引导式临床任务。该联合研究采用整体方法,整合基于知识的声学与文本特征、基于大语言模型的宏观语言描述、基于停顿的声学生物标志物,以及多种神经表示(如LongFormer、ECAPA-TDNN和Trillson嵌入)。将这些特征集与不同分类器结合,生成大量模型,最终选择在训练、开发及各分类性能间取得最佳平衡的模型。结果表明,表现最优的系统由多个互补模型组合而成,依赖于三项临床任务中的声学与文本信息。
原文摘要 · Abstract (English)
This work describes our group's submission to the PROCESS Challenge 2024, with the goal of assessing cognitive decline through spontaneous speech, using three guided clinical tasks. This joint effort followed a holistic approach, encompassing both knowledge-based acoustic and text-based feature sets, as well as LLM-based macrolinguistic descriptors, pause-based acoustic biomarkers, and multiple neural representations (e.g., LongFormer, ECAPA-TDNN, and Trillson embeddings). Combining these feature sets with different classifiers resulted in a large pool of models, from which we selected those that provided the best balance between train, development, and individual class performance. Our results show that our best performing systems correspond to combinations of models that are complementary to each other, relying on acoustic and textual information from all three clinical tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。