arXiv:2603.06961cs.LGcs.AI2026-03

仅用几秒示范数据,就能让四足机器人学会稳定行走。

Learning Quadruped Walking from Seconds of Demonstration

  • 通过分析极限环与神经网络特性,设计新模仿学习方法
  • 几秒演示即可训练出具备鲁棒性的行走策略
  • 适合对少量数据高效训练感兴趣的机器人研究者

四足运动为检验无模型学习是否优于基于模型的控制设计提供了自然场景,其可通过数据模式规避离散接触优化难题和模式切换的组合爆炸问题。本文基于极限环、庞加莱返回映射及神经网络局部数值性质,系统分析了在小数据条件下模仿学习对四足机器人有效的根本原因。由此提出一种新模仿学习方法,通过调节隐空间变化与输出动作变化之间的对齐关系来提升性能。硬件实验表明,仅需数秒示范数据,即可完全离线从零训练出多种具有合理鲁棒性的运动策略。

原文摘要 · Abstract (English)

Quadruped locomotion provides a natural setting for understanding when model-free learning can outperform model-based control design, by exploiting data patterns to bypass the difficulty of optimizing over discrete contacts and the combinatorial explosion of mode changes. We give a principled analysis of why imitation learning with quadrupeds can be inherently effective in a small data regime, based on the structure of its limit cycles, Poincaré return maps, and local numerical properties of neural networks. The understanding motivates a new imitation learning method that regulates the alignment between variations in a latent space and those over the output actions. Hardware experiments confirm that a few seconds of demonstration is sufficient to train various locomotion policies from scratch entirely offline with reasonable robustness.

四足机器人模仿学习小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。