用临床证据引导强化学习,让小模型学会专业精神科推理。
An evidence-guided reinforcement learning method to improve psychiatric reasoning in small language models
- 基于精神科专家策略构建奖励模型,指导小模型优化决策。
- 8B模型超越医学生基准,成为31个模型中表现最佳者。
- 适合医疗AI开发者、精神科研究者及小模型应用落地场景。
隐私与计算限制制约大模型在精神科的应用,而微调小语言模型(SLMs)通常需大量数据与专家标注。我们提出ClinMPO,一种由精神科专家定义的临床精神病学思维策略(CPTS)引导的强化学习框架。ClinMPO使用基于18,569个问答对(来自4,474篇精神科文献)训练的ClinRM奖励模型。在1,737个模型筛选问题上评估了四个Qwen3尺寸。ClinMPO在各规模下均优于基线、监督微调和标准组相对策略优化方法。通过300名高年级医学生的表现建立人类基准,4B模型接近该基准,8B模型则超过基准,并在31个模型及后训练变体中排名第一。ClinMPO在覆盖ICD-11诊断类别与精神科实践能力的两个互补方案中均提升性能。三位盲评临床医生确认其推理质量在CPTS标准上显著改善。结果表明,现有临床证据与专科知识可通过证据引导学习融入医疗AI开发。
原文摘要 · Abstract (English)
Privacy and computational constraints limit the use of large language models in psychiatry, while adapting small language models (SLMs) often requires substantial data and expert annotation. We developed ClinMPO, an evidence-guided reinforcement-learning framework guided by the psychiatrist-defined Clinical Psychiatry Thinking Strategy (CPTS). ClinMPO uses ClinRM, a reward model trained on 18,569 question--answer pairs from 4,474 psychiatry articles. We evaluated four Qwen3 sizes on 1,737 model-screened questions. ClinMPO outperformed Base, supervised fine-tuning and standard group relative policy optimization across scales. From responses by 300 senior pre-licensure medical students, we established the human baseline, a medical-student reference. The 4B model approached this baseline, whereas the 8B model surpassed it and ranked first among 31 models and post-training variants. ClinMPO improved performance across two complementary schemes covering ICD-11 diagnostic categories and psychiatric practice competencies. Blinded assessment by three clinicians showed improved rationale quality across CPTS criteria. These findings highlight how existing clinical evidence and specialist knowledge can be incorporated into the development of medical AI systems through evidence-guided learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。