arXiv:2608.16590cs.RO2026-08

Zetta让机器人在执行中实时自我优化,突破传统学习瓶颈。

Zetta $ζ$: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence

论文配图:Zetta $ζ$: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence
图 1 · 摘自论文原文
  • 构建三层次闭环系统,实现执行时的即时决策与纠错
  • 在LIBERO-Pro和RoboCasa上分别达90.8%和93.6%成功率,推理提速11.1倍
  • 支持零样本迁移与自探索进化,适合追求可靠物理智能的研究者

具身智能体日益用于弥补端到端策略模型的不足。然而,现有方法仍停留在开环执行:固定技能运行,仅在任务结束后反思,无法实时调控物理交互。这是因为当前大模型决策频率远低于环境变化速度。本文提出Zetta,一种闭环具身框架,在保持基础策略冻结的前提下,在线演化代码化运行时评判器与恢复技能。通过三个时间尺度分离的闭环,Zetta实现动作级治理、回放级评判-恢复建议及验证门控的技能更新。结合Z-Infra——一个将代理逻辑与异构执行资源解耦的回放基础设施,Zetta在现有回放预算下达成当前最优表现:LIBERO-Pro成功率达90.8%,RoboCasa达93.6%,推理速度提升11.1倍;成功度随自探索经验持续增长;所学技能可零样本迁移,清晰出现机器人“顿悟”现象。结果表明,闭环框架的自我演化为可靠物理智能提供了可扩展路径。

原文摘要 · Abstract (English)

Embodied agents are increasingly used to close the gap left by end-to-end policy models. Yet the agentic path has not realized closed-loop learning in physical execution: existing harnesses remain largely open-loop, following fixed skills during rollout and reflecting only after an episode completes. Such post-hoc reflection cannot govern execution as it unfolds, because physical interaction requires decisions to track rapidly changing robot-environment states at a frequency beyond today's large agentic models. We present Zetta, a closed-loop embodied harness that evolves code-based runtime critics and recovery skills online while keeping the base policy frozen. Through three timescale-separated loops, Zetta provides action-frequency governance, rollout-level critic-recovery proposal, and validation-gated skill updates. Together with Z-Infra, a rollout infrastructure decoupling agent logic from heterogeneous execution resources, Zetta achieves state-of-the-art success on LIBERO-Pro and RoboCasa under our current rollout budget, reaching 90.8% and 93.6%, with an 11.1x inference speedup; success continues to scale with self-exploration experience; learned skills transfer zero-shot, and clear robotic "Aha Moments" emerge. These results show that closed-loop harness self-evolution opens a scaling path for reliable physical intelligence.

具身智能闭环控制自演化机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。