arXiv:2510.24654cs.CL2025-10被引 4

用强化学习训练会问诊的AI医生,能动态选检查、自适应诊断。

Evolving Interactive Diagnostic Agents in a Virtual Clinical Environment

  • 在虚拟医院环境里通过试错和反馈学诊断策略。
  • 诊断准确率提升11.2%,检查推荐效果提升17.6%。
  • 适合医疗AI研究者与临床决策系统开发者参考。

我们提出一个框架,利用强化学习训练大语言模型作为诊断代理,能够处理多轮交互式问诊,自适应选择检查项目并给出最终诊断。不同于基于静态数据微调的模型,该方法通过动态探索和结果反馈,将患者状态变化映射到最优下一步检查与诊断。主要贡献包括:(i) DiagGym,一个基于电子病历训练的世界模型,构成虚拟临床环境,支持闭环的仿真训练与评估;(ii) DiagAgent,通过端到端多轮强化学习训练,学习优化交互效率与诊断准确性的动态策略;(iii) DiagBench,一个包含2.2K名医验证病例和3.3K份医师评分标准的多中心评测基准,支持过程导向的细粒度评估;(iv) 大规模评估显示,DiagAgent在域内与域外(OOD)设置下均显著优于11个SOTA大模型及2个提示工程代理,在端到端场景中诊断准确率提高11.20%,检查推荐F1得分提升17.58%,并在三个外部中心保持最佳表现。在评分标准评估中,其加权评分比次优模型高7.1%。结果表明,交互式临床环境中学习策略可赋予模型长期管理能力,这是被动训练无法实现的。

原文摘要 · Abstract (English)

We present a framework for training large language models (LLMs) as diagnostic agents with reinforcement learning, enabling them to manage multi-turn interactive diagnostic processes, adaptively select examinations, and commit to final diagnoses. Unlike instruction-tuned models trained on static data, our method acquires diagnostic strategies through dynamic exploration and outcome-based feedback, mapping evolving patient states to the next optimal examination and subsequent diagnosis. Our contributions include: (i) DiagGym, a diagnostics world model trained with electronic health records, serving as a virtual clinical environment to support closed-loop in-silico training and evaluation for interactive diagnosis; (ii) DiagAgent, trained via end-to-end multi-turn RL to learn dynamic diagnostic policies that optimize both interactive effectiveness and final accuracy; (iii) DiagBench, a multi-center diagnostic benchmark designed to evaluate multi-turn diagnostic interaction trajectories. The benchmark comprises 2.2K physician-validated cases sourced from 4 distinct distributions, alongside 3.3K physician-written rubrics for granular process-oriented evaluation. (iv) Extensive evaluations demonstrate DiagAgent's superior performance across both in-domain and out-of-domain (OOD) settings. DiagAgent significantly outperforms 11 SOTA LLMs and 2 prompt-engineered agents. In the end-to-end setting, it delivers a 11.20% increase in diagnostic accuracy and a 17.58% boost in examination recommendation F1 score, while consistently maintaining SOTA performance across all three external centers. Furthermore, in rubric-based evaluations, it surpasses the next-best model by 7.1% in weighted rubric score. These findings indicate that learning policies in interactive clinical environments confers long-term diagnostic management abilities unattainable through passive training.

医疗AI强化学习诊断系统大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。