arXiv:2508.08241cs.RO2025-08被引 254

用扩散模型让人形机器人学会自然流畅的复杂动作并灵活组合应对新任务。

BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion

  • 基于紧凑运动追踪框架与统一扩散模型,单套参数控制多种高难度动作。
  • 实现翻滚、踢腿、冲刺等敏捷动作,且在未训练任务上零样本迁移成功。
  • 适合研究人形机器人控制、动作生成与通用技能泛化的团队参考。

人形机器人的类人形态使其具备人类般的运动敏捷性与多功能性。从人类示范中学习是获取此类能力的可扩展方法。然而,以往工作或产生不自然动作,或依赖特定运动调参才能达到良好自然度。此外,这些方法通常局限于特定动作或目标,缺乏在未见任务中组合多样技能的能力。本文提出BeyondMimic框架,可扩展至多种运动,并在解决未见下游任务时实现无缝技能组合。核心是一个紧凑的运动追踪公式,仅用一套配置和共享超参数,即能掌握包括空中翻滚、旋转踢腿、翻转踢腿和冲刺在内的多种高度敏捷行为,且表现达到当前最优水平。超越简单模仿已有动作,我们引入统一的潜在扩散模型,支持灵活的目标设定、无缝的任务切换及动态动作组合。利用分类器引导(classifier guidance)这一扩散模型特有测试时优化技术,模型可解决训练中未出现的下游任务,如动作补全、手柄遥操作和避障,并实现零样本迁移到真实硬件。该工作通过推动从人类运动中可扩展地学习类人运动技能,开启了人形机器人运动能力的新前沿,实现了超越训练环境的泛化与灵活性。

原文摘要 · Abstract (English)

The human-like form of humanoid robots positions them uniquely to achieve the agility and versatility in motor skills that humans possess. Learning from human demonstrations offers a scalable approach to acquiring these capabilities. However, prior works either produce unnatural motions or rely on motion-specific tuning to achieve satisfactory naturalness. Furthermore, these methods are often motion- or goal-specific, lacking the versatility to compose diverse skills, especially when solving unseen tasks. We present BeyondMimic, a framework that scales to diverse motions and carries the versatility to compose them seamlessly in tackling unseen downstream tasks. At heart, a compact motion-tracking formulation enables mastering a wide range of radically agile behaviors, including aerial cartwheels, spin-kicks, flip-kicks, and sprinting, with a single setup and shared hyperparameters, all while achieving state-of-the-art human-like performance. Moving beyond the mere imitation of existing motions, we propose a unified latent diffusion model that empowers versatile goal specification, seamless task switching, and dynamic composition of these agile behaviors. Leveraging classifier guidance, a diffusion-specific technique for test-time optimization toward novel objectives, our model extends its capability to solve downstream tasks never encountered during training, including motion inpainting, joystick teleoperation, and obstacle avoidance, and transfers these skills zero-shot to real hardware. This work opens new frontiers for humanoid robots by pushing the limits of scalable human-like motor skill acquisition from human motion and advancing seamless motion synthesis that achieves generalization and versatility beyond training setups.

人形机器人扩散模型动作生成零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。