用轻量框架让小模型也能玩转真实世界操作
Guava: An Effective and Universal Harness for Embodied Manipulation

- 设计三要素:循环感知-推理-执行、语义动作抽象、多模态观察
- 仅用2000条仿真轨迹,就让40亿参数模型达到顶尖水平
- 适合想在小模型上实现强泛化操作能力的研究者
基于大规模视觉语言数据训练的语言模型在具身智能体中展现出巨大潜力。通过将模型与外部感知、规划、控制模块结合的工具使用方式,为替代端到端视觉-语言-动作系统提供了新路径。然而,何为高效具身操作框架,以及该框架对各类推理模型的普适性仍不明确。本文提出Guava,一个通过系统探索代理工作流、动作空间和观测空间设计空间而构建的具身工具使用框架。研究识别出三个关键要素:迭代式感知-推理-行动循环、语义动作抽象、多模态观测。为检验这些原则是否适用于小型模型,我们开发了一个端到端训练流程,仅用不到2000条完全在仿真中收集的轨迹,将具身操作能力蒸馏至一个40亿参数的开源模型中。实验结果表明,在仿真和真实环境中的表现可媲美前沿专有模型,并展现出对未见物体、新指令和长程任务的强泛化能力。结果表明,精心设计的框架可作为可扩展、模型无关的具身操作接口,使紧凑的开源模型以极少量训练数据实现强大涌现的具身能力。
原文摘要 · Abstract (English)
Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising alternative to end-to-end vision-language-action systems by combining high-level reasoning with external modules for perception, planning, and control. However, it remains unclear what makes an effective harness for embodied manipulation, and to what extent such a harness can unlock embodied capabilities in a wide range of reasoning models. In this work, we present Guava, a harness framework for embodied tool use developed through systematic exploration of the design space of agent workflows, action spaces, and observation spaces. Our study identifies three key ingredients for effective embodied agents: iterative perception-reasoning-action loops, semantic action abstractions, and multimodal observations. To understand whether these design principles are universal even to small models, we develop an end-to-end training pipeline that distills embodied manipulation capabilities into a 4B open-source model using fewer than 2K trajectories collected entirely in simulation. Experimental results in both simulation and real-world environments show performance comparable to frontier proprietary models while exhibiting strong generalization to unseen objects, novel instructions, and long-horizon tasks. Results suggest that a well-designed harness can serve as a scalable, model-agnostic interface for embodied manipulation, enabling strong emergent embodied capabilities in compact open-source models with minimal training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。