arXiv:2603.22435cs.ROcs.AI2026-03被引 40

用代码当策略,让机器人更智能地完成操作任务

CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation

  • 用可执行代码控制机器人,融合感知与动作原语
  • 12个模型显示:人类设计的抽象越强,表现越好,但依赖人工干预
  • 通过多轮交互和自动技能合成,提升低层指令下的鲁棒性,适合研究具身智能者

Code-as-Policy 探讨可执行代码如何补充数据密集型视觉-语言-动作(VLA)方法,但其在具身操作中作为自主控制器的有效性尚未充分研究。我们提出 CaP-X,一个开源框架,用于系统性研究代码即策略的机器人操作代理。核心是 CaP-Gym,一个交互式环境,代理通过合成并执行组合感知与控制原语的程序来操控机器人。基于此,CaP-Bench 在不同抽象层次、交互方式和感知基础下评估前沿语言与视觉-语言模型。对12个模型的评估显示:性能随人工设计的抽象提升而增强,但去除这些先验后显著下降,暴露对设计者支架的依赖。同时,通过扩展代理运行时计算——包括多轮交互、结构化执行反馈、视觉差分、自动技能合成和集成推理——可显著提升鲁棒性,即使在低层原语上也表现良好。由此推导出 CaP-Agent0,一个无需训练的框架,在模拟和真实机器人上恢复人类级可靠性。我们还引入 CaP-RL,证明使用可验证奖励的强化学习能提高成功率,并实现极小的 sim2real 转移差距。CaP-X 提供了一个原则性且开源的平台,推动具身编码代理的发展。

原文摘要 · Abstract (English)

"Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness as autonomous controllers for embodied manipulation remains underexplored. We present CaP-X, an open-access framework for systematically studying Code-as-Policy agents in robot manipulation. At its core is CaP-Gym, an interactive environment in which agents control robots by synthesizing and executing programs that compose perception and control primitives. Building on this foundation, CaP-Bench evaluates frontier language and vision-language models across varying levels of abstraction, interaction, and perceptual grounding. Across 12 models, CaP-Bench reveals a consistent trend: performance improves with human-crafted abstractions but degrades as these priors are removed, exposing a dependence on designer scaffolding. At the same time, we observe that this gap can be mitigated through scaling agentic test-time computation--through multi-turn interaction, structured execution feedback, visual differencing, automatic skill synthesis, and ensembled reasoning--substantially improves robustness even when agents operate over low-level primitives. These findings allow us to derive CaP-Agent0, a training-free framework that recovers human-level reliability on several manipulation tasks in simulation and on real embodiments. We further introduce CaP-RL, showing reinforcement learning with verifiable rewards improves success rates and transfers from sim2real with minimal gap. Together, CaP-X provides a principled, open-access platform for advancing embodied coding agents.

具身智能代码即策略机器人操作AI代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。