一个能听懂文字、音乐、语音并生成连贯人体动作的通用框架。
OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation
- 用自回归扩散变换器统一建模多种动作生成任务。
- 在28个数据源上构建了最大多模态动捕数据集,支持长时序生成。
- 引入参考动作和渐进式训练,让动作更连贯真实,适合动画与交互应用。
本文提出OmniMotion-X,一种统一序列到序列的自回归扩散变压器框架,用于全身体态动作生成。该框架高效支持文本转动作、音乐跳舞、语音手势、全局时空控制(如动作预测、补全、插帧、关节/轨迹引导生成)等多种多模态任务及其灵活组合。提出以参考动作为新型条件信号,显著提升生成内容、风格与时间动态的一致性。为解决多模态冲突,设计渐进式弱到强混合训练策略。构建了目前最大的统一多模态动作数据集OmniMoCap-X,整合28个公开动捕来源,覆盖10类任务,标准化为SMPL-X格式,30 fps。通过视频渲染与GPT-4o自动生成结构化层级字幕,实现细节与语义的精准标注。大量实验表明,OmniMotion-X在多项任务中超越现有方法,实现高质量、连贯、可控的长时序动作生成。
原文摘要 · Abstract (English)
This paper introduces OmniMotion-X, a versatile multimodal framework for whole-body human motion generation, leveraging an autoregressive diffusion transformer in a unified sequence-to-sequence manner. OmniMotion-X efficiently supports diverse multimodal tasks, including text-to-motion, music-to-dance, speech-to-gesture, and global spatial-temporal control scenarios (e.g., motion prediction, in-betweening, completion, and joint/trajectory-guided synthesis), as well as flexible combinations of these tasks. Specifically, we propose the use of reference motion as a novel conditioning signal, substantially enhancing the consistency of generated content, style, and temporal dynamics crucial for realistic animations. To handle multimodal conflicts, we introduce a progressive weak-to-strong mixed-condition training strategy. To enable high-quality multimodal training, we construct OmniMoCap-X, the largest unified multimodal motion dataset to date, integrating 28 publicly available MoCap sources across 10 distinct tasks, standardized to the SMPL-X format at 30 fps. To ensure detailed and consistent annotations, we render sequences into videos and use GPT-4o to automatically generate structured and hierarchical captions, capturing both low-level actions and high-level semantics. Extensive experimental evaluations confirm that OmniMotion-X significantly surpasses existing methods, demonstrating state-of-the-art performance across multiple multimodal tasks and enabling the interactive generation of realistic, coherent, and controllable long-duration motions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。