arXiv:2608.26800cs.RO2026-08

机器人几分钟内学会多种抛接动作,靠的是边做边学的智能积累。

Rapid On-Robot Learning for Dynamic Manipulation Skills: Robot Juggling

论文配图:Rapid On-Robot Learning for Dynamic Manipulation Skills: Robot Juggling
图 1 · 摘自论文原文
  • 用累积经验建局部模型,保留全局先验,实现持续学习
  • 5分钟内掌握5种三球抛接模式,无需大量试错
  • 兼顾安全与效率,适合真实场景下的机器人自适应训练

我们提出一种在线学习框架,使双臂机器人在物理硬件上仅用几分钟即可直接习得多种抛接模式,即便存在显著的模拟到现实差距。重要启示是:即使模型与真实世界偏差较大,仍可有效辅助学习。因此,本方法的核心理念是:学习应基于机器人已有知识,而非完全替代。通过正则化记忆学习机制,从过往经验中构建局部模型,同时保留全局先验模型以在经验稀疏区域进行外推,从而实现高效稳定的在线学习,避免在庞大行为空间中盲目探索。安全同样关键,我们构建了相互可达集,确保连续抛接动作间可安全过渡,防止任一臂进入需违反关节或执行器限制的状态。结合上述方法,具备多指手和机载视觉的双臂机器人可在不到5分钟的真实交互中,安全地学会并组合五种经典三球抛接模式:交叉式、网球式、半雨刷式、雨刷式和方框式。更广泛地说,该工作为机器人基于不完美先验知识,通过自身真实世界经验持续优化行为提供了新路径。

原文摘要 · Abstract (English)

We present an online learning framework that enables a bimanual robot to acquire diverse juggling patterns directly on physical hardware within minutes, even with a significant sim2real gap. One of the most important lessons from this work is that a model, even when far from reality, can be extremely useful for learning. This motivates a central philosophy of our approach: learning should build upon the robot's current knowledge rather than replace it. Our regularized memory-based learning puts this principle into practice by learning a local model from accumulated experience while retaining the global prior model to extrapolate where experience is sparse. This enables efficient and stable online learning from each new experience without resorting to uninformed exploration over a vast space of possible behaviors. Equally important to continual on-robot learning is safety, allowing the robot to repeatedly practice and improve in the real world. We construct a mutually reachable set that allows safe transitions between successive throws and catches, without driving either arm into a state from which its next action would require violating the robot's joint or actuator limits. Together, these ideas enable a bimanual robot with multi-fingered hands and onboard vision to safely learn and compose five canonical three-ball juggling patterns, including cascade, tennis, half-shower, shower, and box, within less than 5 minutes of real-world interaction. More broadly, this work points toward robots that build upon imperfect prior knowledge and continually refine their behavior through their own real-world experience.

机器人学习在线学习抛接控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。