LLM筛查心理疾病时易忽略症状,尤其当患者功能完好或有社会支持时。
When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening
- 用临床诊断标准构建555例访谈数据集,评估LLM对四类精神障碍的识别能力
- 模型准确率仅0.49–0.86,且对女性、不同种族群体表现不一致
- 当患者有症状但功能正常或有支持系统时,模型常误判为无病
随着心理健康需求超过临床评估供给,可扩展的筛查工具日益重要。大型语言模型(LLMs)可能从患者叙述中识别精神健康风险,但其在不同诊断、人口子组及证据使用模式下的可靠性仍不确定。我们引入一个基于SCID的标准基准,包含555例半结构化体验访谈及其对应诊断标签,涵盖焦虑障碍、重度抑郁障碍、创伤后应激障碍以及当前任何精神障碍。通过零样本任务特定提示,评估五种前沿LLM,并分析假阴性错误是否源于遗漏精神病理证据,或对症状、功能损害与保护性情境线索的权重差异。性能在任务和模型间差异显著,准确率范围为0.49至0.86,马修斯相关系数为0.16至0.38。GPT-4.1 Mini与GPT-5 Mini在疾病特异性准确性上表现最一致。子组分析显示男性抑郁分类准确率高于女性,年龄无稳定趋势,种族间存在轻微非均匀差异。证据整合分析表明,假阴性焦虑与创伤后应激障碍分类常包含明确症状证据,但伴随功能完好、应对能力或社会支持。功能损害证据促使模型向阳性分类倾斜,而保护性情境证据则使其远离阳性判断。这些发现表明,LLM可能辅助大规模精神健康筛查,但在功能完好或有支持背景下忽视症状证据的倾向,需在临床部署前进行严格验证。
原文摘要 · Abstract (English)
As demand for mental health care outpaces clinician-delivered assessment, scalable screening tools are increasingly needed. Large language models (LLMs) may identify psychiatric risk from patient narratives, but their reliability across diagnoses, demographic subgroups, and evidence-use patterns remains uncertain. We introduce a SCID-anchored benchmark of 555 semi-structured experiential interviews paired with diagnostic reference labels for anxiety disorder, major depressive disorder, post-traumatic stress disorder, and any current mental health disorder. Using zero-shot task-specific prompting, we evaluated five state-of-the-art LLMs and examined whether false-negative errors reflected missed psychiatric evidence or differential weighting of symptom, functional-impairment, and protective-context cues. Performance varied across tasks and models, with accuracy ranging from 0.49 to 0.86 and Matthews correlation coefficients from 0.16 to 0.38. GPT-4.1 Mini and GPT-5 Mini showed the most consistent disorder-specific accuracy. Subgroup analyses found higher depression-classification accuracy among male than female participants, no consistent age-related pattern, and modest non-uniform variation across race strata. Evidence-integration analyses showed that false-negative anxiety and PTSD classifications often contained explicit symptom evidence but were accompanied by preserved functioning, coping ability, or social support. Functional-impairment evidence shifted model outputs toward positive classifications, whereas protective-context evidence shifted outputs away. These findings suggest that LLMs may support scalable psychiatric screening, but their tendency to discount symptom evidence in the presence of preserved functioning or protective context requires careful validation before clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。