arXiv:2602.12871cs.CL2026-02

用精神疾病诊断标准构建评测集,检验大模型真能看病吗

MentalBench: A DSM-Grounded Benchmark for Evaluating Psychiatric Diagnostic Capability of Large Language Models

  • 基于DSM-5建立专业知识图谱,生成2.47万例临床案例
  • 模型在清晰信息下答对率高,但面对相似症状时信心失控
  • 适合研究医疗AI可靠性或临床决策支持的开发者

大型语言模型在精神健康评估和临床决策支持中备受关注。然而,现有心理卫生评测多依赖社交媒体数据或对话场景,难以评估模型是否能运用正式诊断标准与鉴别诊断规则。本文提出MentalBench,一个基于DSM-5的评测基准,用于检验大模型在不同临床模糊性下的精神疾病诊断能力。核心是MentalKG——由精神科医生构建并验证的知识图谱,涵盖23种精神障碍的DSM-5诊断标准与鉴别规则。利用该图谱生成24,750个合成临床案例,系统控制信息完整度与诊断复杂度,实现基于DSM的评估。实验表明,尽管先进大模型在无噪声查询中表现良好,但在区分症状重叠的疾病时无法准确校准自身置信度。这提示当前大模型作为精神科辅助工具的可靠性存疑,亟需更贴近真实诊疗挑战的评估体系。

原文摘要 · Abstract (English)

Large language models (LLMs) have attracted growing interest as supportive tools for psychiatric assessment and clinical decision support. However, existing mental health benchmarks largely rely on social media data or supportive dialogue settings, limiting their ability to assess whether models can apply formal diagnostic criteria and differential diagnostic rules. In this paper, we introduce MentalBench, a benchmark for evaluating whether LLMs can make DSM-grounded psychiatric diagnostic decisions under varying levels of clinical ambiguity. At the core of MentalBench is MentalKG, a psychiatrist-built and validated knowledge graph encoding DSM-5 diagnostic criteria and differential diagnostic rules for 23 psychiatric disorders. Using MentalKG as an expert-curated logical backbone, we generate 24,750 synthetic clinical cases that systematically vary in information completeness and diagnostic complexity, enabling DSM-grounded evaluation. Our experiments show that although state-of-the-art LLMs perform well on noise-free queries that probe DSM-5 knowledge, they struggle to calibrate their confidence when distinguishing between disorders with overlapping symptoms. These findings raise concerns about the reliability of LLMs as psychiatric decision-support tools and highlight the need for more evaluation that reflects the diverse challenges in real-world psychiatric diagnosis.

精神健康诊断评估DSM-5大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。