首个面向中文精神科问诊的多智能体评测基准,测试大模型诊断能力。
LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis
- 构建1.6万条真实分布的模拟问诊对话,覆盖12类精神疾病。
- 模型在共病识别上准确率仅43%,动态问诊比静态评估差。
- 问得有逻辑不等于诊断准,适合研究中文医疗大模型的学者。
精神障碍全球高发,但精神科医生短缺与基于访谈诊断的主观性,严重阻碍了及时、一致的心理健康评估。当前人工智能辅助精神科诊断进展受限,缺乏同时具备真实患者模拟、临床验证诊断标签及支持多轮动态问诊的评测基准。本文提出LingxiDiagBench,一个大规模多智能体评测框架,用于评估大模型在中文精神科问诊与诊断中的表现。核心为LingxiDiag-16K数据集,包含16,000条与电子病历对齐的合成问诊对话,涵盖12类ICD-10精神疾病,还原真实临床人口学与诊断分布。在主流大模型上的实验发现:(1)二分类抑郁-焦虑任务准确率达92.3%,但共病识别下降至43.0%,12类鉴别诊断仅为28.5%;(2)动态问诊表现普遍低于静态评估,表明无效信息获取策略显著影响诊断推理;(3)由大模型作为评判者评估的问诊质量与诊断准确率相关性仅中等,说明结构良好提问未必带来正确诊断。我们已开源LingxiDiag-16K与完整评测框架,以支持可复现研究(https://github.com/Lingxi-mental-health/LingxiDiagBench)。
原文摘要 · Abstract (English)
Mental disorders are highly prevalent worldwide, but the shortage of psychiatrists and the inherent subjectivity of interview-based diagnosis create substantial barriers to timely and consistent mental-health assessment. Progress in AI-assisted psychiatric diagnosis is constrained by the absence of benchmarks that simultaneously provide realistic patient simulation, clinician-verified diagnostic labels, and support for dynamic multi-turn consultation. We present LingxiDiagBench, a large-scale multi-agent benchmark that evaluates LLMs on both static diagnostic inference and dynamic multi-turn psychiatric consultation in Chinese. At its core is LingxiDiag-16K, a dataset of 16,000 EMR-aligned synthetic consultation dialogues designed to reproduce real clinical demographic and diagnostic distributions across 12 ICD-10 psychiatric categories. Through extensive experiments across state-of-the-art LLMs, we establish key findings: (1) although LLMs achieve high accuracy on binary depression--anxiety classification (up to 92.3%), performance deteriorates substantially for depression--anxiety comorbidity recognition (43.0%) and 12-way differential diagnosis (28.5%); (2) dynamic consultation often underperforms static evaluation, indicating that ineffective information-gathering strategies significantly impair downstream diagnostic reasoning; (3) consultation quality assessed by LLM-as-a-Judge shows only moderate correlation with diagnostic accuracy, suggesting that well-structured questioning alone does not ensure correct diagnostic decisions. We release LingxiDiag-16K and the full evaluation framework to support reproducible research at https://github.com/Lingxi-mental-health/LingxiDiagBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。