用推理模型提升语音痴呆分类,但需避免幻觉误导。
Do Multimodal Large Language Models Need Reasoning to Classify Dementia from Speech?

- 设计适配器框架DeTAiL,利用推理模型内部表示而非文本理由。
- 在两个数据集上均超越基线和依赖文本理由的方法。
- 适合关注医疗诊断可解释性与模型鲁棒性的研究者。
多模态大语言模型(MLLMs)为提升语音记录中自动痴呆分类(ADC)系统的准确性、可迁移性和可解释性提供了新思路。然而,其推理能力是否有助于ADC仍不明确,如何有效利用也尚不清楚。本文对推理型MLLMs在ADC中的应用进行了细致评估,发现直接依赖文本推理过程会导致诊断幻觉和不一致,性能反而低于无LLM的基线。为此,我们提出基于适配器的DeTAiL框架,通过挖掘推理型MLLM的内部表征实现更优的痴呆分类。在两个测试格式与标签粒度不同的痴呆数据集上,DeTAiL始终优于强基线及依赖文本理由的方法。代码与演示将在论文接受后发布。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have emerged as a promising approach for improving the accuracy, transferability, and explainability of automatic dementia classification (ADC) systems from voice recordings. Yet it remains unclear whether their reasoning capabilities are beneficial for ADC, and how such capabilities should be leveraged. In this paper, we conduct a careful evaluation of reasoning MLLMs for ADC and show that naive strategies, such as relying on text-based rationales, can lead to hallucinated and inconsistent rationales for diagnosis and yield inferior ADC performance compared with LLM-free baselines. To overcome this limitation, we propose \textbf{De}mentia \textbf{T}hinker with Nonlinear \textbf{A}daptor and Re\textbf{i}nforcement \textbf{L}earning (DeTAiL), an adaptor-based framework that exploits the internal representations of reasoning MLLMs for improved dementia classification. Across two dementia datasets with distinct test formats and label granularities, DeTAiL consistently outperforms strong baselines and methods that rely on text-based rationales. Code and demo will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。