arXiv:2608.22797cs.AI2026-08

专为精神科患者教育打造的AI模型,回答药物问题更完整,但医生更偏爱通用大模型。

Performance of a domain-specific large language model in answering patient questions in psychiatry

  • 用权威医疗资料微调专属精神科LLM,提升临床准确性
  • 在完整性上优于ChatGPT,但医生评分更倾向通用模型
  • 适合精神科患者教育场景,助力安全医患沟通

本研究评估了专用于精神科领域的大语言模型(LLM)MIND在回答患者关于舍曲林问题上的表现,对比其与ChatGPT和OpenEvidence的差异。我们使用两种方法:(1)基于准确度、清晰度、完整性、细微性、安全性及转诊适当性等维度的计算机评分;(2)由10名持证精神科医师进行主观评分。结果显示,按评分标准,MIND在所有维度均显著优于其他模型(p<0.001)。但在医师评分中,ChatGPT的准确性略高于MIND(p=0.021, r=0.073),MIND在完整性上显著更优(p<0.001, r=0.160),安全性无显著差异(p=0.955, r=0.002)。多数医师(57.6%)偏好ChatGPT生成的回答,而42.4%选择MIND(p=0.003)。结论:尽管MIND在多数指标上表现更优,但医生仍更倾向于通用大模型的输出。该模型为构建安全可靠的临床辅助系统迈出重要一步。

原文摘要 · Abstract (English)

Background This study was designed to evaluate whether a domain-specific large language model (LLM) trained exclusively on patient education resources can answer questions about psychiatric medications, in a manner superior to LLM chatbots. We developed an LLM ("MIND") fine-tuned for clinical fidelity, trained on patient education resources from authoritative medical organizations. Methods We compared the responses of MIND, ChatGPT, and OpenEvidence to patient questions about escitalopram, using two methods: (1) computer analysis according to a rubric measuring accuracy, clarity, completeness, nuance, safety, and referral appropriateness; (2) ratings from N=10 board-licensed psychiatrists on similar metrics. Results When rated by rubric, MIND was rated highest in all domains (p<0.001). When rated by psychiatrists, ChatGPT was rated accurate more often than MIND with a negligible effect size (p=0.021, r=0.073); MIND was rated complete more often than ChatGPT with a small effect size (p<0.001, r=0.160); and MIND and ChatGPT were rated safe with the same frequency (p=0.955, r=0.002). The majority of psychiatrists preferred the responses generated by ChatGPT (57.6%) compared to MIND (42.4%, p=0.003). Conclusions MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time. However, despite MIND's ability to provide more complete responses, psychiatrists preferred ChatGPT's responses. MIND represents a step towards building safe LLM systems to enhance patient education in psychiatry.

精神科AI大模型患者教育

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。