arXiv:2608.07418cs.AIcs.CL2026-08

用强化学习训练AI模拟医生问诊,提升诊断准确率和临床决策能力。

ResidencyRL: Reinforcement Learning in Simulated Clinical Environments

论文配图:ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
图 1 · 摘自论文原文
  • 通过多轮对话模拟真实问诊,用强化学习训练医疗AI代理。
  • 诊断准确率提升7.0%,红灯信号遗漏减少31%,专家更倾向使用该AI。
  • 适用于需要复杂推理的临床场景,适合医学AI研发与教育研究者。

在医学教育中,医生通过数年住院培训,在数千次患者互动中将理论知识转化为临床技能,经历多样反馈并逐步获得更大自主权。临床推理高度依赖于问诊过程——医生需通过对话获取病史、修正诊断假设,并在不确定性下制定管理方案。尽管大语言模型(LLMs)在静态医学基准上表现优异,但对完整临床决策序列的优化方法仍不成熟。本文提出ResidencyRL,一种基于强化学习(RL)的方法,通过模拟多轮临床问诊(每条轨迹最多60轮对话和8次工具调用)训练临床人工智能(AI)代理。ResidencyRL将策略代理与具备复杂对抗行为的大语言模型模拟器配对,采用结构化奖励函数,涵盖诊断准确性、管理质量、沟通、文档记录和安全。在保留评估中,该代理在对抗条件下诊断准确率提升7.0%(88.0% vs. 81.0%),红灯信号遗漏率降低31%,有效缓解过早闭合问题。盲评专家临床医生在对比中偏好该代理达87.6%。其操作能力可迁移至未见基准:在AMIE多访视基准的全部六个临床维度上优于基础模型,并在AgentClinic和CRAFT-MD上持续呈现正向改进。结果表明,通过模拟中的多轮强化学习,可有效习得序列化临床决策能力,获得稳健且可泛化的性能,为迈向临床精通提供路径。未来需在真实工作流中开展前瞻性验证以确立临床实用性。

原文摘要 · Abstract (English)

In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy. Much of clinical reasoning relies on the patient encounter, a dialogue in which a clinician elicits history, refines diagnostic hypotheses, and decides management under uncertainty. While large language models (LLMs) excel on static medical benchmarks, methods to optimize the full sequence of clinical decisions remain underdeveloped. We present ResidencyRL, a reinforcement learning (RL) method for training clinical artificial intelligence (AI) agents through simulated multi-turn clinical encounters (up to 60 dialogue turns and 8 tool calls per trajectory). ResidencyRL pairs the policy agent with LLM simulators capable of complex, adversarial behaviors, training against a structured reward aligned to diagnostic accuracy, management quality, communication, documentation, and safety. On held-out evaluations, the ResidencyRL agent improves diagnostic accuracy by 7.0% under adversarial conditions (88.0% vs. 81.0%) and reduces missed red flag rates by 31%, demonstrating rigorous mitigation of premature closure. Blinded expert clinicians validated these gains, preferring the trained agent in 87.6% of side-by-side comparisons. The procedural competencies transfer to unseen benchmarks: the agent outperforms the base model across all six clinical axes of the AMIE multi-visit benchmark, and shows consistent directional improvements on AgentClinic and CRAFT-MD. Our findings demonstrate that sequential clinical decision-making can be effectively learned through multi-turn RL in simulation, yielding robust, generalizable capabilities, paving the way towards clinical mastery. Prospective validation with real-world workflows remains necessary to establish clinical utility.

强化学习临床决策大模型应用医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。