让视觉语言动作模型用现成工具,提升机器人任务泛化能力。
Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use

- 将现成工具注入视觉语言动作模型,降低动作空间复杂度。
- 在3万条轨迹数据上训练,成功率比主流方法高20%。
- 适合需要低数据依赖、可扩展工具的机器人应用。
本文将端到端视觉-语言-动作(VLA)模型与代理式工具使用结合,提出工具使用型代理机器人(ART)。ART是一种工具注入框架,可微调任意VLA模型,利用现成工具模块完成低层视觉、高层功能识别和具身增强。相比具有连续动作空间的原始VLA模型,ART通过工具使用降低了动作空间复杂度,不仅提升了跨任务泛化能力,还减少了数据依赖。为验证该框架优势,我们构建了包含3万条工具使用轨迹和操作示范的数据集,远小于基线方法所用规模。同时设计了在复杂环境中长轨迹工具推理的训练方案。实验表明,ART在模拟与真实场景任务中成功率比主流基线高出20%,如在新视角下黑暗环境中进行拾取放置。实证结果表明代理式方法的优势:模块化工具使用实现更高效训练、轻量级部署及新工具可扩展集成。该设计增强了鲁棒性、适应性和可拓展性,为复杂现实场景中VLA系统的实际部署铺平道路。
原文摘要 · Abstract (English)
This paper integrates end-to-end Visual-Language-Action (VLA) models with agentic tool-use to propose Agentic Robot with Tool-use (ART). ART is a tool-injection framework that tunes any VLA model to leverage off-the-shelf tool modules for low-level vision, high-level affordance, and embodiment enhancement. Compared to vanilla VLA models with a whole continuous action solution space, ART reduces the complexity of the action solution space through tool-use, which not only improves generalizability across different tasks but also reduces data dependency. To demonstrate the advantages (high generalizability and low data dependency) of this framework, we first built a dataset of 30K tool-use trajectories and action demonstrations, which is much smaller than those used by baseline methods. We then designed a training regimen for long-trajectory tool-use reasoning in challenging environments. Experiments show that ART achieves a 20% higher success rate than mainstream baselines on simulation and real-world tasks, such as pick-and-place in the dark at novel viewpoints. Empirical results highlight the benefits of an agent-based approach: modular tool utilization enables more efficient training, lightweight deployment, and scalable integration of new tools. This design fosters robustness, adaptability, and extensibility, paving the way for the practical deployment of VLA systems in complex real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。