arXiv:2607.02983cs.AI2026-07被引 1

让大模型像医生一样主动找证据诊断,提升医疗推理能力。

Reinforcement Learning for Evidence-Seeking Diagnostic Reasoning with Large Language Models

论文配图:Reinforcement Learning for Evidence-Seeking Diagnostic Reasoning with Large Language Models
图 1 · 摘自论文原文
  • 用强化学习驱动大模型主动寻找诊断线索,模拟临床问诊过程。
  • 在多个数据集上表现接近更大模型,且生成的检查建议更符合医学逻辑。
  • 适合医疗AI研究者、临床辅助系统开发者参考。

近期以推理为核心的大型语言模型(LLMs)虽取得显著进展,但大多采用被动推理模式,假设信息完全可用。而真实临床智能是迭代式的调查过程,需主动获取证据。为此,本文将医学诊断形式化为一个迭代证据获取任务,利用可验证奖励的强化学习(RLVR)在闭环环境中激发模型内在推理能力,并设计新奖励机制以保证诊断准确性和检查一致性。为支持该框架,提出基于检索增强生成的检查模拟器(RAGES),作为高保真临床决策代理,提供知识驱动的真实后续证据。实验证明,该框架使大模型从被动应答者转变为自主助手,在多种数据集上性能与更大、具备推理增强的基线模型相当;同时,RAGES生成的临床反馈在生物学合理性上优于普通大模型。

原文摘要 · Abstract (English)

Recent reasoning-centric Large Language Models (LLMs) have made significant strides, yet they predominantly operate on a passive-inference pattern that assumes complete information. In contrast, real-world clinical intelligence is inherently an iterative investigative process requiring strategic evidence acquisition. To bridge this gap, we formalize medical diagnosis as an Iterative Evidence-Seeking Task. We leverage Reinforcement Learning with Verifiable Rewards (RLVR) to elicit intrinsic reasoning within a closed-loop environment, guided by a novel suite of rewards that enforce diagnostic precision and examination consistency. To facilitate this, we introduce the Retrieval-Augmented Generation-based Examination Simulator (RAGES), a high-fidelity clinical oracle that provides realistic, knowledge-grounded follow-up evidence. Empirical results across diverse datasets demonstrate that our framework enables LLMs to transition from passive responders to autonomous assistants. Notably, our model demonstrates comparable performance to larger and reasoning-enhanced baselines, while RAGES proves superior to vanilla LLMs in generating biologically plausible clinical feedback.

医疗AI强化学习推理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。