arXiv:2505.20289cs.CV2025-05被引 34

让视觉智能体学会自主选工具,靠试错提升推理能力

VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection

  • 用强化学习让代理自动探索和选择视觉工具
  • 在多个测试集上显著优于无需训练的基线方法
  • 适合需要灵活适应新任务的视觉推理系统

我们提出VisTA,一种新的强化学习框架,使视觉智能体能够基于实际表现,动态探索、选择并组合多样工具库中的工具。现有工具增强型推理方法或依赖无训练提示,或需大规模微调;两者均缺乏主动工具探索,通常假设工具种类有限,且微调方法还需大量人工标注。相比之下,VisTA通过端到端强化学习,利用任务结果作为反馈信号,迭代优化针对查询的复杂工具选择策略。通过群体相对策略优化(GRPO),该框架使代理能在无需显式推理监督的情况下,自主发现有效的工具选择路径。在ChartQA、Geometry3K和BlindTest基准上的实验表明,VisTA在分布外样本上相较于无训练基线实现显著性能提升。结果凸显其增强泛化能力、自适应利用多样化工具的潜力,为可灵活演进的视觉推理系统铺平道路。

原文摘要 · Abstract (English)

We introduce VisTA, a new reinforcement learning framework that empowers visual agents to dynamically explore, select, and combine tools from a diverse library based on empirical performance. Existing methods for tool-augmented reasoning either rely on training-free prompting or large-scale fine-tuning; both lack active tool exploration and typically assume limited tool diversity, and fine-tuning methods additionally demand extensive human supervision. In contrast, VisTA leverages end-to-end reinforcement learning to iteratively refine sophisticated, query-specific tool selection strategies, using task outcomes as feedback signals. Through Group Relative Policy Optimization (GRPO), our framework enables an agent to autonomously discover effective tool-selection pathways without requiring explicit reasoning supervision. Experiments on the ChartQA, Geometry3K, and BlindTest benchmarks demonstrate that VisTA achieves substantial performance gains over training-free baselines, especially on out-of-distribution examples. These results highlight VisTA's ability to enhance generalization, adaptively utilize diverse tools, and pave the way for flexible, experience-driven visual reasoning systems.

视觉推理强化学习工具选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。