arXiv:2607.08257cs.AI2026-07

构建虚拟精神科问诊环境,评估大模型临床能力。

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

论文配图:MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters
图 1 · 摘自论文原文
  • 基于1193份病历构建标准化患者,模拟完整诊疗流程。
  • 大模型在客观能力上比医生低37.28个百分点,评估是短板。
  • 专设五类评分器,适合研究临床对话与医疗AI的团队使用。

大型语言模型在孤立的精神科任务中表现优异,但现有基准很少模拟完整的临床问诊过程。我们提出《MentalHospital》,一个用于评估基于大模型的精神科临床问诊的虚拟环境。该环境实现主观访谈、客观检查、诊断评估和治疗计划(S.O.A.P.)全流程,基于1,193份去标识化精神科电子病历(EHR)构建,覆盖所有主要ICD-11类别和76种疾病。每次问诊通过双轨评估:一是与病历参考进行客观对比,二是对临床过程质量进行主观打分。为规模化专家评判,我们开发了MentalEval,包含五类领域专用评估器,分别评估沟通共情、问诊专业性、病历质量、诊断严谨性和治疗适宜性,采用基于评分标准的SFT和专家引导的DPO训练。22名临床医师调查反馈显示,MentalHospital具有较高的临床真实感(3.88/5),MentalEval与专家评分平均达到0.944的QWK一致性。基准测试表明,即使最强的LLM在客观精神科能力上仍落后于临床医生37.28个百分点,其中精神状态评估是主要瓶颈。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters. We introduce $\textbf{MentalHospital}$, a virtual evaluation environment for LLM-based psychiatric clinical encounters. MentalHospital instantiates the Subjective Interviewing, Objective Examination, Diagnostic Assessment, and Treatment Planning (S.O.A.P.) workflow, using skill-augmented standardized patients constructed from 1,193 de-identified psychiatric electronic health record (EHR) cases spanning all major ICD-11 categories and 76 disorders. Each encounter is assessed through a dual-track protocol that combines objective comparison against EHR-derived references with subjective assessment of clinical process quality. To scale specialist judgment, we develop $\textbf{MentalEval}$, five domain-specific evaluators covering communication empathy, interviewing professionalism, clinical-note quality, diagnostic rigor, and treatment appropriateness, trained with rubric-grounded SFT and expert-guided DPO. Survey responses from 22 clinicians support MentalHospital's clinical fidelity (3.88/5), while MentalEval achieves strong expert alignment with an average QWK of 0.944. Benchmarking shows that even the strongest LLM trails clinicians by 37.28 percentage points in objective psychiatric competence, with mental status assessment as a key bottleneck.

精神科AI临床评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。