让视觉模型学会组合使用工具并动态调整策略,提升智能体的视觉推理能力。
VC-Tooler: Learning Compositional and Adaptive Visual Tool Use

- 通过分层合成构建轨迹库,覆盖单工具定位、多工具组合与多样接口场景
- 两阶段训练:监督冷启动建立基础能力,强化学习优化准确性与效率
- 在多个基准上表现领先,支持复杂工具环境下的推理迁移
智能体多模态推理通过视觉工具交互主动获取和修正视觉证据,拓展了被动图像理解。有效视觉工具使用需具备三项能力:在视觉上下文中定位工具调用、跨步骤组合工具、根据工具返回结果自适应调整推理。然而现有方法多聚焦于固定工具空间内的定位与固定调用模式,对组合与自适应能力关注不足。本文提出VC-Tooler,将视觉工具使用建模为可组合且自适应的能力。首先,通过分层合成流程构建轨迹库,涵盖三个能力层级:单工具定位、多工具组合、多样工具上下文与界面。随后进行两阶段训练:监督冷启动建立基础能力,强化学习推动准确、高效、上下文感知的视觉工具使用。VC-Tooler在通用与智能体基准上均达到开源模型最优表现,包括在V*上达95.8%,在VTC-Bench上达35.3%,并在推理时面对更丰富的工具设置展现出良好泛化能力。
原文摘要 · Abstract (English)
Agentic multimodal reasoning extends passive image understanding by allowing VLMs to actively acquire and refine visual evidence through visual tool interactions. Effective visual tool use requires three capabilities: grounding tool calls in visual context, composing tools across multiple steps, and adapting reasoning to tool-returned observations. However, existing approaches largely focus on grounding within fixed tool spaces and rigid invocation patterns, leaving composition and adaptation insufficiently addressed. We present VC-Tooler, which learns visual tool use as a compositional and adaptive capability. To this end, we first build a trajectory bank through a hierarchical synthesis pipeline covering three capability levels: single-tool grounding, multi-tool composition, and diverse tool contexts and interfaces. We then train the model in two stages: a supervised cold start that establishes these capabilities, followed by reinforcement learning that encourages accurate, efficient, and context-aware visual tool use. VC-Tooler achieves state-of-the-art performance among open-source models on both general-purpose and agentic benchmarks, including $95.8\%$ on V* and $35.3\%$ on VTC-Bench, and shows promising transfer under richer tool settings at inference time. Project page: https://w1zheng.github.io/VC-Tooler
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。