构建首个基于权威精神科教材的多任务评测基准,验证大模型临床能力。
PsychiatryBench: A Multi-Task Benchmark for LLMs in Psychiatry
- 基于专家认证的精神科教材与案例集构建评测数据
- 覆盖11类临床任务,共5188个标注问题,涵盖诊断与管理全流程
- 发现主流大模型在随访与治疗规划中存在安全与一致性短板
大型语言模型(LLMs)在提升精神科诊疗效率方面潜力巨大,涵盖诊断准确性、临床记录自动化及治疗支持等。然而,现有评估资源多依赖小型临床访谈语料、社交媒体文本或合成对话,缺乏临床真实性且无法反映诊断推理的复杂性。本文提出PsychiatryBench,一个完全基于权威专家验证的精神科教科书与案例集构建的多任务评测基准。该基准包含11项问答任务,涵盖诊断推理、治疗计划、纵向随访、管理规划、临床策略、序列病例分析以及选择题/扩展匹配题,总计5,188个专家标注样本。我们评估了包括Google Gemini、DeepSeek、Sonnet 4.5和GPT-5在内的前沿大模型,以及MedGemma等开源医学模型,采用传统指标与“大模型作为评判者”的相似度评分框架。结果揭示模型在多轮随访与管理任务中存在显著临床一致性与安全性缺口,凸显需针对性调优与更严谨的评估范式。PsychiatryBench提供模块化、可扩展平台,助力精神健康领域大模型性能提升。
原文摘要 · Abstract (English)
Large language models (LLMs) offer significant potential in enhancing psychiatric practice, from improving diagnostic accuracy to streamlining clinical documentation and therapeutic support. However, existing evaluation resources heavily rely on small clinical interview corpora, social media posts, or synthetic dialogues, which limits their clinical validity and fails to capture the full complexity of diagnostic reasoning. In this work, we introduce PsychiatryBench, a rigorously curated benchmark grounded exclusively in authoritative, expert-validated psychiatric textbooks and casebooks. PsychiatryBench comprises eleven distinct question-answering tasks ranging from diagnostic reasoning and treatment planning to longitudinal follow-up, management planning, clinical approach, sequential case analysis, and multiple-choice/extended matching formats totaling 5,188 expert-annotated items. {\color{red}We evaluate a diverse set of frontier LLMs (including Google Gemini, DeepSeek, Sonnet 4.5, and GPT 5) alongside leading open-source medical models such as MedGemma using both conventional metrics and an "LLM-as-judge" similarity scoring framework. Our results reveal substantial gaps in clinical consistency and safety, particularly in multi-turn follow-up and management tasks, underscoring the need for specialized model tuning and more robust evaluation paradigms. PsychiatryBench offers a modular, extensible platform for benchmarking and improving LLM performance in mental health applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。