让机器人自主优化操作策略,实现真实世界下的高效自进化。
ENPIRE: Agentic Robot Policy Self-Improvement in the Real World

- 构建闭环系统,自动重置场景、执行策略、验证结果并迭代优化。
- 在复杂操作任务中实现99%成功率,支持多机器人并行训练。
- 适合希望减少人工干预、推动机器人自主进化的研究者和工程师。
真实世界中实现灵巧机器人操作高度依赖人工监督与算法工程,成为通用物理智能发展的核心瓶颈。尽管新兴编码代理可生成代码自动化算法搜索,但其成功仍局限于数字环境。我们提出,自动化机器人研究的缺失抽象是可重复的现实世界策略改进反馈循环:重置场景、执行策略、验证结果、优化下一轮。为此,我们引入ENPIRE框架,通过四个核心模块实现这一物理反馈流程:环境模块(EN)自动重置与验证,策略优化模块(PI)启动策略改进,滚动执行模块(R)利用一个或多个实体机器人并行评估策略,进化模块(E)则由编码代理分析日志、查阅文献、改进训练基础设施与算法代码以应对失败模式。该闭环系统将真实世界操作学习转化为可控优化过程,显著降低人力投入,并支持训练方案与代理变体的公平消融实验。基于ENPIRE,前沿编码代理可自主训练出在复杂灵巧操作任务(如整理图钉盒、系扎尼龙绳、工具使用)中达到99%成功率的策略,当部署多代理团队于机器人集群时,优化速度进一步提升。结果表明,编码代理可为真实世界机器人自主进步提供一条可行且可扩展的路径。
原文摘要 · Abstract (English)
Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intelligence. Although emerging coding agents can generate code to automate algorithm search, their successes remain largely confined in digital environments. We conjecture that the missing abstraction to automate robotics research is a repeatable feedback loop for real-world policy improvement: reset the scene, execute a policy, verify the outcome, and refine the next iteration. To bridge this gap, we introduce ENPIRE, a harness framework for coding agents that instantiates this physical feedback routine with four core modules: an Environment module (EN) for automatic reset and verification, a Policy Improvement module (PI) that launches policy refinement, a Rollout module (R) to evaluate policies with one or multiple physical robots operating in parallel, and an Evolution module (E) in which coding agents analyze logs, consult literature, improve training infrastructure and algorithm code to address failure modes. This closed-loop system transforms real-world manipulation learning into a controllable optimization procedure, minimizing human effort while allowing fair ablations across training recipe and agent variants. Powered by ENPIRE, frontier coding agents can autonomously train a policy to achieve a 99% success rate on challenging, dexterous manipulation tasks, such as organizing a pin box, fastening a zip tie, and tool use, a process that further accelerates when we dispatch an agent team on a robot fleet. Our results suggest a practical and scalable path toward deploying coding agents to autonomously advancing robotics in the physical world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。