arXiv:2606.05143cs.RO2026-06

提出可恢复性约束的课程学习,让机器人在物理世界训练更稳健。

HORIZON: Recoverability-Governed Curriculum for Physical-Domain Scaling

论文配图:HORIZON: Recoverability-Governed Curriculum for Physical-Domain Scaling
图 1 · 摘自论文原文
  • 基于可恢复性设计分阶段物理域扩展课程
  • 实验证明盲目扩域会降低鲁棒性,需有序推进
  • 适合做具身智能与机器人强化学习的研究者

机器人策略的鲁棒性提升不仅需要更广的随机化,还需保持训练过程中物理经验的可学习性。我们研究了策略何时能从更复杂的物理环境获益,发现可恢复性是在线策略训练中物理域扩展的核心约束。在线训练中,新动力学只有在足够接近当前策略以生成可纠正的数据时才有效,否则会导致无法恢复的失败。以四足机器人行走为具身泛化基准,我们提出HORIZON——一种基于检查点的前沿课程方法,仅在当前策略可恢复边界内扩展物理域。该方法通过回滚和边界精炼控制每一步扩展,将固定随机化转变为持续的物理域增长过程。实验揭示三个规律:第一,直接全域扩展在不同物理维度上不均衡,常需分步有序进行;第二,域组合具有非单调性,超出紧凑核心后可能稀释可恢复联合样本,降低整体鲁棒性;第三,离线蒸馏孤立专家无法替代在线课程产生的联合交互。这些结果将物理域泛化视为具身控制中的持续生长问题,以可恢复性作为在线扩展的组织原则。

原文摘要 · Abstract (English)

Scaling robust robot policies requires more than broader randomization, because physical-domain experience must remain organized and learnable throughout training. We study when a policy can benefit from harder physics and identify recoverability as a central constraint in on-policy physical-domain scaling. In on-policy training, new dynamics are useful only insofar as they remain close enough to the current policy to generate corrective on-policy data, rather than collapsing rollouts into unrecoverable failures. Using quadruped locomotion as a physically demanding benchmark for embodied generalization, we introduce HORIZON, a checkpointed frontier curriculum that expands physical domains only within the current policy's recoverable boundary. HORIZON uses rollback and boundary refinement to govern each expansion step, turning fixed randomization into a continual process of physical-domain growth. Experiments reveal three regularities of physical-domain expansion. First, direct domain widening is uneven across physical axes and often unlearnable without staged ordering. Second, domain composition is non-monotonic, and adding more domains beyond a compact core can dilute recoverable joint samples and reduce overall robustness. Third, offline distillation of isolated experts cannot substitute for the joint interaction generated by on-policy curriculum. Together, these results frame physical-domain generalization as a continual growth problem for embodied control, with recoverability as the organizing principle for on-policy expansion.

机器人学习强化学习课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。