arXiv:2607.26809cs.RO2026-07

机器人零样本自进化,通过反复练习自主提升操作能力

Practice Makes Policies: Bootstrapping and Consolidating Robotic Capabilities from Zero Human Demonstrations

论文配图:Practice Makes Policies: Bootstrapping and Consolidating Robotic Capabilities from Zero Human Demonstrations
图 1 · 摘自论文原文
  • 构建自驱式框架,从零人类示范中自发积累经验
  • 实现技能快速复用与闭环策略固化,跨任务表现更稳健
  • 适合追求自适应机器人系统的研究者和开发者

通用机器人操作需在开放世界环境中执行多样化任务,并持续提升技能。现有系统多以静态方式学习特定任务或场景下的能力,缺乏通过物理交互自适应演进的机制。受人类通过反复练习形成肌肉记忆的启发,高级操作能力依赖于让机器人将交互经验逐步转化为高效操作能力的自主演化机制。为此,我们提出HERO——一种无需人类示范即可实现能力自进化的分层具身智能体。HERO将启发式推理、范例复用与反射式执行统一整合为协同框架,使机器人能够自主启动操作经验积累,通过经验迁移快速生成可复用行为,并将重复交互逐步固化为高效的闭环视觉运动策略。通过紧密耦合自主数据采集与任务执行,HERO根据经验积累阶段和执行需求动态扩展并调度操作能力。大量实验表明,HERO显著减少机器人数据收集中的手动干预,同时在多种任务下均表现出鲁棒操作能力,为自进化机器人系统提供了可行路径。

原文摘要 · Abstract (English)

General-purpose robotic manipulation requires robots to perform diverse tasks in open-world environments while improving their skills over time. Despite recent progress in robotic manipulation, existing systems still primarily acquire manipulation skills in a static manner, where capabilities are learned for specific tasks or settings rather than adaptively evolving through physical interaction. Resembling how repeated practice enables humans to develop muscle memory, advanced manipulation proficiency requires an autonomous capability evolution mechanism that allows robots to progressively transform interaction experiences into increasingly effective manipulation abilities. To this end, we propose HERO, a self-improving hierarchical embodied agent that enables autonomous capability evolution from zero human demonstrations. HERO organizes heuristic reasoning, exemplar reuse, and reflexive execution into a unified orchestration framework, allowing robots to autonomously bootstrap manipulation experience, rapidly accumulate reusable behaviors through experience transfer, and progressively consolidate recurring interactions into efficient closed-loop visuomotor policies. By tightly coupling autonomous data collection with task execution, HERO continuously expands and dynamically schedules manipulation capabilities according to different stages of experience accumulation and execution requirements. Extensive experiments demonstrate that HERO substantially reduces human intervention during robotic data collection while achieving robust manipulation across diverse tasks, providing a promising path toward self-improving robotic systems.

机器人自进化零样本策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。