arXiv:2503.01903cs.CLcs.AI2025-03被引 8

构建心理临床评估基准,测试大模型实际应用效果。

PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice

  • 设计心理临床专用评测框架PsychBench,融合真实诊疗需求
  • 16个大模型测试显示现有模型仍不足以独立决策
  • 对初级医生帮助显著,可提升效率与临床质量

大语言模型(LLMs)为缓解精神科医疗资源短缺和诊断一致性低的问题提供了潜在解决方案。然而,当前缺乏一个全面、专业的基准框架来评估LLMs在真实精神科临床环境中的表现,制约了面向精神科应用的专用大模型发展。针对这一空白,我们结合精神科临床需求与真实数据,提出了PsychBench评测系统,用于评估LLMs在精神科临床场景中的实用性。我们对16个主流大模型进行了全面定量评估,探究提示工程、思维链推理、输入文本长度及领域知识微调对模型性能的影响。通过详细错误分析,识别出现有模型的优势与局限,并提出改进方向。随后,开展了包含60名不同资历精神科医生的临床读者研究,进一步探索现有大模型作为辅助工具的实际价值。量化评估与读者研究均表明:尽管现有模型潜力显著,但尚不能作为独立决策工具;作为辅助工具,其对初级医生具有明显支持作用,能有效提升工作效率与整体临床质量。为推动该领域研究,我们将公开数据集与评测框架,助力大模型在精神科临床实践中的落地应用。

原文摘要 · Abstract (English)

The advent of Large Language Models (LLMs) offers potential solutions to address problems such as shortage of medical resources and low diagnostic consistency in psychiatric clinical practice. Despite this potential, a robust and comprehensive benchmarking framework to assess the efficacy of LLMs in authentic psychiatric clinical environments is absent. This has impeded the advancement of specialized LLMs tailored to psychiatric applications. In response to this gap, by incorporating clinical demands in psychiatry and clinical data, we proposed a benchmarking system, PsychBench, to evaluate the practical performance of LLMs in psychiatric clinical settings. We conducted a comprehensive quantitative evaluation of 16 LLMs using PsychBench, and investigated the impact of prompt design, chain-of-thought reasoning, input text length, and domain-specific knowledge fine-tuning on model performance. Through detailed error analysis, we identified strengths and potential limitations of the existing models and suggested directions for improvement. Subsequently, a clinical reader study involving 60 psychiatrists of varying seniority was conducted to further explore the practical benefits of existing LLMs as supportive tools for psychiatrists of varying seniority. Through the quantitative and reader evaluation, we show that while existing models demonstrate significant potential, they are not yet adequate as decision-making tools in psychiatric clinical practice. The reader study further indicates that, as an auxiliary tool, LLM could provide particularly notable support for junior psychiatrists, effectively enhancing their work efficiency and overall clinical quality. To promote research in this area, we will make the dataset and evaluation framework publicly available, with the hope of advancing the application of LLMs in psychiatric clinical settings.

精神科大模型评测临床辅助心理学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。