对比大模型与心理医生诊断人格障碍,发现模型在自传式叙述中表现差异显著。
Patterns vs. Patients: Evaluating LLMs against Mental Health Professionals on Personality Disorder Diagnosis through First-Person Narratives
- 用波兰语自述文本对比大模型与专业医生诊断能力
- 大模型对边缘型障碍识别准确率超医生,但对自恋型几乎误诊
- 模型依赖模式分析,医生更关注患者自我体验与时间感知
随着大语言模型在精神健康自评中的广泛应用,其解读患者第一人称叙述的能力受到关注。本研究通过深度案例比较,评估了当前最先进大模型与心理健康专业人士在基于波兰语第一人称自传体叙述下对边缘型(BPD)和自恋型(NPD)人格障碍的诊断表现。结果显示,表现最佳的Gemini Pro模型总体诊断得分(65.48%)比人类专家平均得分(43.57%)高出21.91个百分点。尽管两者在识别BPD方面均表现优异(F1=83.4 vs. 80.0),但模型对NPD的识别严重不足(F1=6.7 vs. 50.0),可能因‘自恋’一词带有价值判断而回避。定性分析显示,模型给出自信且详尽的模式化解释,而人类专家则保持简洁谨慎,更关注患者的自我感与时间经验。研究揭示,尽管大模型能处理复杂临床叙事,但其可靠性与偏见问题仍不容忽视。
原文摘要 · Abstract (English)
Growing reliance on LLMs for psychiatric self-assessment raises questions about their ability to interpret qualitative patient narratives. This depth over breadth case study directly compares state-of-the-art LLMs and mental health professionals in assessing Borderline (BPD) and Narcissistic (NPD) Personality Disorders based on Polish-language first-person autobiographical accounts. Within our sample, the overall diagnostic scores of the top-performing Gemini Pro models (65.48%) were 21.91 percentage points higher than the average scores of the human professionals (43.57%). While both models and human experts excelled at identifying BPD (F1 = 83.4 & F1 = 80.0, respectively), models severely underdiagnosed NPD (F1 = 6.7 vs. 50.0), showing a potential reluctance toward the value-laden term "narcissism." Qualitatively, models provided confident, elaborate justifications focused on patterns and formal categories, while human experts remained concise and cautious, emphasizing the patients' sense of self and temporal experience. Our findings demonstrate that while LLMs might be competent at interpreting complex first-person clinical data, their outputs still carry critical reliability and bias issues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。