让大模型像医生一样主动找证据诊断,提升医疗推理能力。
Reinforcement Learning for Evidence-Seeking Diagnostic Reasoning with Large Language Models

- 用强化学习驱动大模型主动寻找诊断线索,模拟临床问诊过程。
- 在多个数据集上表现接近更大模型,且生成的检查建议更符合医学逻辑。
- 适合医疗AI研究者、临床辅助系统开发者参考。
近期以推理为核心的大型语言模型(LLMs)虽取得显著进展,但大多采用被动推理模式,假设信息完全可用。而真实临床智能是迭代式的调查过程,需主动获取证据。为此,本文将医学诊断形式化为一个迭代证据获取任务,利用可验证奖励的强化学习(RLVR)在闭环环境中激发模型内在推理能力,并设计新奖励机制以保证诊断准确性和检查一致性。为支持该框架,提出基于检索增强生成的检查模拟器(RAGES),作为高保真临床决策代理,提供知识驱动的真实后续证据。实验证明,该框架使大模型从被动应答者转变为自主助手,在多种数据集上性能与更大、具备推理增强的基线模型相当;同时,RAGES生成的临床反馈在生物学合理性上优于普通大模型。
原文摘要 · Abstract (English)
Recent reasoning-centric Large Language Models (LLMs) have made significant strides, yet they predominantly operate on a passive-inference pattern that assumes complete information. In contrast, real-world clinical intelligence is inherently an iterative investigative process requiring strategic evidence acquisition. To bridge this gap, we formalize medical diagnosis as an Iterative Evidence-Seeking Task. We leverage Reinforcement Learning with Verifiable Rewards (RLVR) to elicit intrinsic reasoning within a closed-loop environment, guided by a novel suite of rewards that enforce diagnostic precision and examination consistency. To facilitate this, we introduce the Retrieval-Augmented Generation-based Examination Simulator (RAGES), a high-fidelity clinical oracle that provides realistic, knowledge-grounded follow-up evidence. Empirical results across diverse datasets demonstrate that our framework enables LLMs to transition from passive responders to autonomous assistants. Notably, our model demonstrates comparable performance to larger and reasoning-enhanced baselines, while RAGES proves superior to vanilla LLMs in generating biologically plausible clinical feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。