用视频生成经验提升动作生成的泛化能力
The Quest for Generalizable Motion Generation: Data, Model, and Evaluation
- 从视频生成迁移数据、模型与评估框架,打通跨领域知识
- 构建22.8万样本数据集,融合动捕、视频与合成数据
- 提出轻量模型和分层评测基准,适合追求泛化的研究者
尽管在标准基准上3D人体动作生成(MoGen)取得进展,现有文本到动作模型仍面临泛化能力的根本瓶颈。相比之下,视频生成(ViGen)在建模人类行为方面展现出显著泛化能力,为MoGen提供了可借鉴的思路。受此启发,我们提出一个系统性框架,从数据、建模到评估三个关键层面,将ViGen知识迁移至MoGen。首先,引入ViMoGen-228K,一个包含22.8万条高质量动作样本的大规模数据集,融合高保真光学动捕数据、网络视频中的语义标注动作及先进ViGen模型生成的合成样本,包含文本-动作对与文本-视频-动作三元组,大幅扩展语义多样性。其次,提出基于流匹配的扩散变换器ViMoGen,通过门控多模态条件统一动捕数据与ViGen模型先验;为提升效率,进一步设计去视频依赖的轻量版ViMoGen-light,仍保持强泛化能力。最后,构建MBench,一个分层评测基准,用于细粒度评估动作质量、提示一致性和泛化能力。大量实验表明,该框架在自动与人工评估中均显著优于现有方法。代码、数据与基准将公开。主页:https://motrixlab.github.io/2026_iclr_vimogen。
原文摘要 · Abstract (English)
Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing text-to-motion models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generative fields, most notably video generation (ViGen), have demonstrated remarkable generalization in modeling human behaviors, highlighting transferable insights that MoGen can leverage. Motivated by this observation, we present a comprehensive framework that systematically transfers knowledge from ViGen to MoGen across three key pillars: data, modeling, and evaluation. First, we introduce ViMoGen-228K, a large-scale dataset comprising 228,000 high-quality motion samples that integrates high-fidelity optical MoCap data with semantically annotated motions from web videos and synthesized samples generated by state-of-the-art ViGen models. The dataset includes both text-motion pairs and text-video-motion triplets, substantially expanding semantic diversity. Second, we propose ViMoGen, a flow-matching-based diffusion transformer that unifies priors from MoCap data and ViGen models through gated multimodal conditioning. To enhance efficiency, we further develop ViMoGen-light, a distilled variant that eliminates video generation dependencies while preserving strong generalization. Finally, we present MBench, a hierarchical benchmark designed for fine-grained evaluation across motion quality, prompt fidelity, and generalization ability. Extensive experiments show that our framework significantly outperforms existing approaches in both automatic and human evaluations. The code, data, and benchmark will be made publicly available. Homepage: https://motrixlab.github.io/2026_iclr_vimogen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。