用音频大模型零样本检测认知障碍,无需标注也能跨语言通用。
Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM
- 用提示词引导音频大模型直接判断语音是否反映认知异常。
- 在英美和多语言数据集上表现接近有监督方法,跨任务一致。
- 适合无标注数据或需快速适配新语言/场景的研究者使用。
认知障碍(CI)日益成为公共健康问题,早期检测对干预至关重要。语音作为非侵入性且易获取的生物标志物受到关注。传统方法依赖人工标注提取声学与语言特征的监督模型,泛化能力受限。本文首次提出基于Qwen2-Audio AudioLLM的零样本语音认知障碍检测方法,该模型可处理音频与文本输入。通过设计提示词指令,引导模型将语音样本分类为正常或认知障碍。在英语和多语言两个数据集上评估,涵盖不同认知评估任务。结果表明,零样本音频大模型性能接近监督方法,且在语言、任务与数据集间表现出良好泛化性与一致性。
原文摘要 · Abstract (English)
Cognitive impairment (CI) is of growing public health concern, and early detection is vital for effective intervention. Speech has gained attention as a non-invasive and easily collectible biomarker for assessing cognitive decline. Traditional CI detection methods typically rely on supervised models trained on acoustic and linguistic features extracted from speech, which often require manual annotation and may not generalise well across datasets and languages. In this work, we propose the first zero-shot speech-based CI detection method using the Qwen2- Audio AudioLLM, a model capable of processing both audio and text inputs. By designing prompt-based instructions, we guide the model in classifying speech samples as indicative of normal cognition or cognitive impairment. We evaluate our approach on two datasets: one in English and another multilingual, spanning different cognitive assessment tasks. Our results show that the zero-shot AudioLLM approach achieves performance comparable to supervised methods and exhibits promising generalizability and consistency across languages, tasks, and datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。