让医疗视觉模型无需人工标注就能自主推理,提升诊断可靠性。
MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning
- 用模型不确定度引导视觉探索,自动定位关键图像区域。
- 通过多轮推理结果一致性的自动生成伪标签,驱动学习。
- 在多个公开数据集上超越现有模型,适合临床安全场景应用。
医学视觉-语言模型在复杂临床任务中潜力巨大,但其推理能力常受限于仅依赖文本的范式,难以基于视觉证据进行判断,不仅影响细粒度分析表现,还可能导致安全敏感场景下的视觉幻觉。为此,我们提出MedVR,一种无需人工标注中间步骤的强化学习框架,实现医疗VLM的无标注视觉推理。核心创新包括:熵引导的视觉再定位(EVR),利用模型不确定性指导探索;基于共识的信用分配(CCA),从推理轨迹的一致性中提炼伪监督信号。该方法在多个公开医疗VQA基准上达到领先性能,显著优于现有模型。通过直接与视觉证据交互进行推理,MedVR提升了模型鲁棒性与可解释性,助力医疗AI的临床落地。
原文摘要 · Abstract (English)
Medical Vision-Language Models (VLMs) hold immense promise for complex clinical tasks, but their reasoning capabilities are often constrained by text-only paradigms that fail to ground inferences in visual evidence. This limitation not only curtails performance on tasks requiring fine-grained visual analysis but also introduces risks of visual hallucination in safety-critical applications. Thus, we introduce MedVR, a novel reinforcement learning framework that enables annotation-free visual reasoning for medical VLMs. Its core innovation lies in two synergistic mechanisms: Entropy-guided Visual Regrounding (EVR) uses model uncertainty to direct exploration, while Consensus-based Credit Assignment (CCA) distills pseudo-supervision from rollout agreement. Without any human annotations for intermediate steps, MedVR achieves state-of-the-art performance on diverse public medical VQA benchmarks, significantly outperforming existing models. By learning to reason directly with visual evidence, MedVR promotes the robustness and transparency essential for accelerating the clinical deployment of medical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。