用AI编程代理零样本解决推T块任务,效率远超人类示范学习。
Revisiting the "Push-T" Robot Manipulation Task with Agentic Robotics

- AI代理通过自动生成仿真代码,无需演示数据自主优化推物策略。
- 在2D仿真中达成100%成功率,用时比最优扩散模型少46%。
- 可扩展至全字母推写字母任务,支持真实机械臂3D仿真与视觉反馈。
Push-T是基于人类示范学习操纵策略的经典基准任务:机器人需以单点接触将T形块推至目标姿态。本文在智能体机器人新范式下重访该任务,让大语言模型编码代理(Claude Code with Fable 5)生成无需示范数据的算法解法。实验显示,该代理成功定位2D仿真环境,通过仿真实验学习推动物理机制,并迭代优化,最终实现100%成功率,且所需步数比使用200次人类示范训练的最优扩散模型减少46%。此外,该代理还通过自生成课程扩展至从Push-A到Push-Z的全字母任务,为Franka和UR5机械臂生成3D跨平台仿真代码,并支持带视觉反馈的仿真运行。视频、策略与详细信息将公开发布。
原文摘要 · Abstract (English)
Push-T is an iconic benchmark for learning manipulation policies from human demonstrations. The robot must use a single point of contact to push a T-shaped block into a target pose. In this short paper, we revisit the Push-T task in the context of emerging advances in Agentic Robotics where an LLM coding agent -- Claude Code with Fable 5 -- is prompted to create an algorithmic solution that does not require any demonstration data. We study how effective the agentic coding loop can solve the Push-T task, and compare the resulting code as policy with the visuomotor imitation learning policy. Results suggest that the agent found the 2D gym simulation online, and used sim experiments to learn push mechanics, iteratively optimizing to achieve 100% success rate using 46% fewer steps than the best diffusion policy trained with 200 human demonstrations. The coding agent also solve extensions from T to the full alphabet (Push-A to Push-Z) using a self generated curriculum and generated simulation code for the Franka and UR5 robot arms in 3D cross-embodiment simulations with visual feedback. Videos, policies and details will be posted online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。