arXiv:2605.29488cs.CVcs.AI2026-05被引 1

用大规模多模态数据训练,实现任意条件组合下的高质量动作生成。

AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling

论文配图:AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
图 1 · 摘自论文原文
  • 用残差分层量化编码动作,结合掩码建模的Transformer统一处理多模态输入。
  • 在超过5000小时的数据上训练,支持文本、语音、音乐等任意组合控制。
  • 适合需要灵活动作控制的虚拟人、游戏动画和机器人研发人员。

条件化人体动作生成是计算机视觉与机器人领域的重要挑战。尽管已有显著进展,现有方法常受限于固定模态配置和任务特定架构,导致跨模态交互及多模态合成的扩展规律研究不足。主要瓶颈在于大规模对齐动作数据稀缺,制约了不同控制信号间的泛化能力。本文提出OmniHuMo,一个包含超过5000小时动作数据和320万条序列的大规模高质量多模态数据集,具备精确对齐的文本、语音、音乐和轨迹等标注。基于该数据集,我们提出AnyMo,一种统一的多模态框架,结合基于残差分层量化(Residual FSQ)的动作分词器与可扩展的掩码建模Transformer,实现任意模态组合下的高质量动作合成。大量实验表明,AnyMo在保持高保真度的同时,可灵活控制空间与风格属性。

原文摘要 · Abstract (English)

Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configurations and task-specific architectures, leaving cross-modal interactions and the scaling laws of multimodal-conditioned synthesis largely underexplored. A key bottleneck is the scarcity of large-scale modality-aligned motion data, limiting generalization across diverse control signals. In this work, we introduce OmniHuMo, a large-scale, high-quality dataset comprising over 5,000 hours of motion and 3.2 million sequences with precisely aligned multimodal annotations (e.g., text, speech, music, and trajectory). Leveraging OmniHuMo, we propose AnyMo, a unified multimodal framework combining a Residual FSQ-based motion tokenizer with a scalable masked modeling transformer, enabling high-quality motion synthesis under arbitrary modality combinations. Extensive experiments show that AnyMo achieves high-fidelity synthesis while offering flexible control over both spatial and stylistic attributes.

动作生成多模态掩码建模大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。