让视觉语言模型学会用工具一步步思考图像,提升复杂视觉推理能力。
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
- 构建统一接口的VISTA-Gym环境,支持多步视觉工具交互与强化学习训练。
- VISTA-R1模型在11个基准上性能领先同类模型9.5%至18.7%,显著提升工具协同能力。
- 适合研究视觉智能体、多模态推理与强化学习融合的学者参考。
尽管近期视觉语言模型(VLMs)具备强大的图像理解能力,但其通过多步视觉交互进行推理的能力仍受限。本文提出VISTA-Gym,一个可扩展的训练环境,用于激励VLMs发展工具集成的视觉推理能力。该环境统一了来自13个数据集的7类真实世界多模态推理任务,提供标准化的视觉工具接口(如定位、解析)、可执行的交互循环、可验证的反馈信号及高效的轨迹记录,支持大规模视觉智能体强化学习。现有VLMs虽在纯文本推理上表现良好,但在工具选择、调用与协调方面仍存不足。借助VISTA-Gym,我们训练出VISTA-R1,通过多轮轨迹采样与端到端强化学习,实现工具使用与智能体推理的交替。在11个公开的推理密集型VQA基准上的实验表明,VISTA-R1-8B在性能上比同类规模的最先进模型高出9.51%至18.72%,证明VISTA-Gym是激发VLMs工具集成推理能力的有效训练平台。
原文摘要 · Abstract (English)
While recent vision-language models (VLMs) demonstrate strong image understanding, their ability to "think with images", i.e., to reason through multi-step visual interactions, remains limited. We introduce VISTA-Gym, a scalable training environment for incentivizing tool-integrated visual reasoning capabilities in VLMs. VISTA-Gym unifies diverse real-world multimodal reasoning tasks (7 tasks from 13 datasets in total) with a standardized interface for visual tools (e.g., grounding, parsing), executable interaction loops, verifiable feedback signals, and efficient trajectory logging, enabling visual agentic reinforcement learning at scale. While recent VLMs exhibit strong text-only reasoning, both proprietary and open-source models still struggle with tool selection, invocation, and coordination. With VISTA-Gym, we train VISTA-R1 to interleave tool-use with agentic reasoning via multi-turn trajectory sampling and end-to-end reinforcement learning. Extensive experiments across 11 public reasoning-intensive VQA benchmarks show that VISTA-R1-8B outperforms state-of-the-art baselines with similar sizes by 9.51%-18.72%, demonstrating VISTA-Gym as an effective training ground to unlock the tool-integrated reasoning capabilities for VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。