用诊断语言预测抑郁得分的模型存在数据污染,真实效果被夸大。
"Mirror" Language AI Models of Depression are Criterion-Contaminated
- 用诊断问答文本预测抑郁分,导致结果虚高
- 非诊断语言也能预测抑郁,效果仍显著
- 建议改用外部语言提升模型临床可信度
近期研究显示,基于语言的抑郁评分预测模型表现接近完美(R² = .70),但这些“镜像”模型依赖于抑郁评估本身的语言回答来预测评估分数,存在准则污染问题。本研究对比了“镜像”模型与“非镜像”模型,后者使用其他外部语言(如生活史访谈)预测抑郁分。110名参与者完成了结构化诊断访谈(镜像条件)和生活史访谈(非镜像条件)。大语言模型根据两种语言生成预测。结果表明,镜像模型预测近乎完美;而即使在非镜像条件下,模型预测效应量在心理学中仍属较大。此外,两类模型与问卷式抑郁症状的相关性相近,提示镜像模型存在偏差。主题建模显示不同模型类型间主题结构差异明显。随着语言模型在心理评估中的发展,采用非镜像方法可能提升其有效性和临床实用性。
原文摘要 · Abstract (English)
Recent studies show near-perfect language-based predictions of depression scores (R2 = .70), but these "Mirror" models rely on language responses directly from depression assessments to predict depression assessment scores. These methods suffer from criterion contamination that inflate prediction estimates. We compare "Mirror" models to "Non-Mirror" models, which use other external language to predict depression scores. 110 participants completed both structured diagnostic (Mirror condition) and life history (Non-Mirror condition) interviews. LLMs were prompted to predict diagnostic depression scores. As expected, Mirror models were near-perfect. However, Non-Mirror models also displayed prediction sizes considered large in psychology. Further, both Mirror and Non-Mirror predictions correlated with other questionnaire-based depression symptoms at similar sizes, suggesting bias in Mirror models. Topic modeling revealed different theme structures across model types. As language models for depression continue to evolve, incorporating Non-Mirror approaches may support more valid and clinically useful language-based AI applications in psychological assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。