用临床推理框架提升大模型精神疾病诊断能力
MentalSeek-Dx: Towards Progressive Hypothetico-Deductive Reasoning for Real-world Psychiatric Diagnosis
- 构建真实临床场景的诊断基准,引导模型学习假设演绎推理
- 仅140亿参数即达顶尖水平,显著优于现有模型在细粒度诊断上的表现
- 适合医疗AI研究者和精神科临床辅助系统开发者参考
精神健康障碍是全球日益严峻的公共卫生挑战。尽管大语言模型在精神评估中展现潜力,但其临床应用受限于缺乏生态有效性与细粒度诊断监督的基准。为此,我们提出首个面向真实临床环境的疾病级别精神诊断基准——MentalDx Bench,包含712份去标识电子病历,由持证精神科医师按ICD-11标准标注,覆盖16类诊断中的76种障碍。对18个LLM的评估显示:粗粒度分类表现良好,但在疾病级别诊断上系统性失败,暴露出模式匹配与临床假设演绎推理之间的范式错位。针对此问题,我们提出MentalSeek-Dx,一种通过监督轨迹构建与课程强化学习训练的医学专用大模型,内化临床推理过程。在MentalDx Bench上的实验表明,MentalSeek-Dx仅用140亿参数即达到最先进性能,建立了可信赖的精神病诊断临床基础框架。
原文摘要 · Abstract (English)
Mental health disorders represent a burgeoning global public health challenge. While Large Language Models (LLMs) have demonstrated potential in psychiatric assessment, their clinical utility is severely constrained by benchmarks that lack ecological validity and fine-grained diagnostic supervision. To bridge this gap, we introduce \textbf{MentalDx Bench}, the first benchmark dedicated to disorder-level psychiatric diagnosis within real-world clinical settings. Comprising 712 de-identified electronic health records annotated by board-certified psychiatrists under ICD-11 guidelines, the benchmark covers 76 disorders across 16 diagnostic categories. Evaluation of 18 LLMs reveals a critical \textit{paradigm misalignment}: strong performance at coarse diagnostic categorization contrasts with systematic failure at disorder-level diagnosis, underscoring a gap between pattern-based modeling and clinical hypothetico-deductive reasoning. In response, we propose \textbf{MentalSeek-Dx}, a medical-specialized LLM trained to internalize this clinical reasoning process through supervised trajectory construction and curriculum-based reinforcement learning. Experiments on MentalDx Bench demonstrate that MentalSeek-Dx achieves state-of-the-art (SOTA) performance with only 14B parameters, establishing a clinically grounded framework for reliable psychiatric diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。