让机器人能听懂语言、音频和动作,生成自然全身动作。
OMG: Omni-Modal Motion Generation for Generalist Humanoid Control

- 用扩散模型融合语言、音频和动作输入生成运动
- 在多个新模态下实现高效适应与泛化
- 为通用人形机器人控制提供可扩展的智能核心
近年来人形机器人全身控制取得显著进展,但现有方法仍局限于少数技能策略且依赖复杂奖励设计,或难以拓展至新输入模态的运动追踪器。本文认为通用人形控制的关键在于构建一个可扩展的“大脑”——能够处理多种条件输入的模块,叠加于反应式运动追踪“小脑”之上,模仿生物运动系统的层级结构。实现该愿景面临两大挑战:获取大规模高质量数据以支持通用控制,以及赋予生成器对组合性、可扩展多模态输入的条件建模能力。为此,我们提出OMG,通过精心设计的数据清洗、过滤与标注流程,结合基于扩散模型的运动生成主干网络,支持语言、音频及人类参考动作的多模态输入。大量实验验证了OMG作为全模态全身控制器的先进性能,展现出优秀的模型缩放行为与对新分布、新模态的高效适应能力,标志着向人形机器人基础模型迈出了实质性一步。
原文摘要 · Abstract (English)
Humanoid whole-body control has made significant progress in recent years, yet existing approaches remain limited to few-skill policies with heavy reward engineering, or motion trackers that are difficult to extend to new input modalities. We argue that the key to general-purpose humanoid control is to build a scalable brain, a module capable of reasoning with diverse conditioning modalities, atop a reactive motion tracking cerebellum, mirroring the hierarchical structure of biological motor systems. Two challenges arise in realizing this vision: acquiring a vast amount of high-quality data to achieve general purpose control, and equipping the generator with the capability to condition on compositional, extensible multi-modal inputs. We present OMG, which addresses these challenges with a meticulous data curation, filtering and labeling pipeline, as well as a diffusion-based motion generation backbone that conditions on language, audio, and human reference motions. Extensive experiments validate OMG as an omni-modal whole-body controller exhibiting state-of-the-art performance, model scaling behavior and efficient adaptation to new distributions and modalities, marking a concrete step toward foundation models for humanoid robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。