arXiv:2605.09634cs.CL2026-05

评测大模型在心理筛查中表现,发现三类可靠性问题。

Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness

论文配图:Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness
图 1 · 摘自论文原文
  • 用零样本语音预测焦虑抑郁分数,测试模型一致性与抗语音识别误差能力。
  • 部分模型在语音识别错误率10%时,评分一致性从0.82骤降至0.36。
  • 模型给出的结论依据不靠谱,关键词支持度不足,影响临床可信度。

大语言模型可零样本地通过语音估计医院焦虑抑郁量表(HADS)得分,但临床应用需保障三方面可靠性:模型内一致性、语音识别(ASR)鲁棒性以及证据忠实性。我们对三个模型(Phi-4、Gemma-2-9B、Llama-3.1-8B)在111名英语使用者上进行评估,使用真实转录文本和三种Whisper ASR变体(Large、Medium、Small),每组模型-条件组合运行三次。结果表明:(i) Phi-4与Gemma-2-9B具有优异的模型内一致性(ICC > 0.89),在ASR扰动下几乎无退化;(ii) Llama-3.1-8B对ASR敏感,当误识率(WER)达10%时,一致性从0.82降至0.36;(iii) 稳定模型在ASR干扰下仍保持较高预测有效性;(iv) Phi-4与Gemma-2-9B的关键词依附性超过93%,而Llama-3.1-8B下降至77%-81%。不同模型间关键词一致率远低于评分一致性,揭示评分与证据间的脱节,影响临床解释性。

原文摘要 · Abstract (English)

LLMs can estimate Hospital Anxiety and Depression Scale (HADS) scores from speech in a zero-shot manner, but clinical deployment requires reliability across three dimensions: intra-model consistency, ASR robustness, and evidence faithfulness. We evaluate three LLMs (Phi-4, Gemma-2-9B, and Llama-3.1-8B) on 111 English-speaking participants using ground-truth transcripts and three Whisper ASR variants (Large, Medium, Small), with three independent runs per model-condition pair. We find that (i) Phi-4 and Gemma-2-9B achieve excellent intra-model consistency (ICC > 0.89) with minimal degradation under ASR; (ii) Llama-3.1-8B shows ASR-fragile consistency, with ICC dropping from 0.82 to 0.36 at 10% WER; (iii) predictive validity is largely preserved under ASR for robust models; and (iv) keyword groundedness exceeds 93% for Phi-4 and Gemma-2-9B but falls to 77-81% for Llama-3.1-8B. Inter-model keyword agreement is far lower than score-level agreement, revealing a score-evidence dissociation with implications for clinical interpretability.

心理筛查大模型评估语音分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。