arXiv:2601.07107cs.CVcs.AI2026-01被引 3

让医学影像模型学会用工具一步步思考,提升诊断推理能力。

MEDVISTAGYM: A Scalable Training Environment for Thinking with Medical Images via Tool-Integrated Reinforcement Learning

  • 构建可扩展交互环境,训练模型自主选择并使用工具进行多步推理。
  • 在6个医学问答基准上,性能超越同类模型19.10%至24.21%。
  • 适合需要复杂医学图像分析的AI研究者与临床辅助系统开发者。

视觉语言模型(VLMs)在通用图像理解中表现优异,但在医学图像的多步推理任务中表现不佳,主要受限于静态视觉嵌入和单次推理机制,难以反复验证或修正视觉证据。尽管引入工具的推理路径具有潜力,但开源VLM缺乏有效训练基础设施以学习工具选择、调用与协同。我们提出MedVistaGym,一个可扩展的交互式训练环境,激励模型通过工具集成实现医学图像分析中的主动推理。该环境使模型能够判断何时调用工具、定位关键图像区域,并将单个或多个子图像证据整合进交错的多模态推理流程中,支持代理式训练。基于MedVistaGym,我们训练了MedVistaGym-R1,通过轨迹采样与端到端强化学习实现工具使用与代理推理的融合。在六个医学视觉问答基准上,MedVistaGym-R1-8B相比同规模工具增强基线模型提升19.10%至24.21%,证明结构化代理训练而非单纯工具接入,才是实现高效医学图像工具融合推理的关键。

原文摘要 · Abstract (English)

Vision language models (VLMs) achieve strong performance on general image understanding but struggle to think with medical images, especially when performing multi-step reasoning through iterative visual interaction. Medical VLMs often rely on static visual embeddings and single-pass inference, preventing models from re-examining, verifying, or refining visual evidence during reasoning. While tool-integrated reasoning offers a promising path forward, open-source VLMs lack the training infrastructure to learn effective tool selection, invocation, and coordination in multi-modal medical reasoning. We introduce MedVistaGym, a scalable and interactive training environment that incentivizes tool-integrated visual reasoning for medical image analysis. MedVistaGym equips VLMs to determine when and which tools to invoke, localize task-relevant image regions, and integrate single or multiple sub-image evidence into interleaved multimodal reasoning within a unified, executable interface for agentic training. Using MedVistaGym, we train MedVistaGym-R1 to interleave tool use with agentic reasoning through trajectory sampling and end-to-end reinforcement learning. Across six medical VQA benchmarks, MedVistaGym-R1-8B exceeds comparably sized tool-augmented baselines by 19.10% to 24.21%, demonstrating that structured agentic training--not tool access alone--unlocks effective tool-integrated reasoning for medical image analysis.

医学图像工具推理强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。