让动作生成理解人类意图并结合视觉信息,提升精准度与可控性。
MoGIC: Boosting Motion Generation via Intention Understanding and Visual Context
- 融合意图建模与视觉先验,统一生成动作与理解行为动机。
- 在HumanML3D和Mo440H上分别降低38.6%和34.6%的FID值。
- 支持文本、视觉条件及意图预测,适合需要高精度动作控制的场景。
现有文本驱动的动作生成方法多将语言与动作视为双向映射,但难以捕捉行为执行的因果逻辑与人类意图。缺乏视觉锚定也限制了生成的精确性与个性化,因语言无法指定细粒度时空细节。本文提出MoGIC,一个整合意图建模与视觉先验的统一框架。通过联合优化多模态条件下的动作生成与意图预测,MoGIC挖掘潜在人类目标,利用视觉先验增强生成质量,并展现多功能生成能力。我们进一步设计自适应作用范围的注意力混合机制,实现条件标记与动作子序列间的有效局部对齐。为支持该范式,我们构建了440小时的基准数据集Mo440H,涵盖21个高质量运动数据集。实验表明,微调后MoGIC在HumanML3D上降低38.6% FID,Mo440H上降低34.6% FID;在动作描述任务中优于基于LLM的方法(轻量文本头);并支持意图预测与视觉条件生成,推动可控制动作合成与意图理解的发展。代码已开源。
原文摘要 · Abstract (English)
Existing text-driven motion generation methods often treat synthesis as a bidirectional mapping between language and motion, but remain limited in capturing the causal logic of action execution and the human intentions that drive behavior. The absence of visual grounding further restricts precision and personalization, as language alone cannot specify fine-grained spatiotemporal details. We propose MoGIC, a unified framework that integrates intention modeling and visual priors into multimodal motion synthesis. By jointly optimizing multimodal-conditioned motion generation and intention prediction, MoGIC uncovers latent human goals, leverages visual priors to enhance generation, and exhibits versatile multimodal generative capability. We further introduce a mixture-of-attention mechanism with adaptive scope to enable effective local alignment between conditional tokens and motion subsequences. To support this paradigm, we curate Mo440H, a 440-hour benchmark from 21 high-quality motion datasets. Experiments show that after finetuning, MoGIC reduces FID by 38.6\% on HumanML3D and 34.6\% on Mo440H, surpasses LLM-based methods in motion captioning with a lightweight text head, and further enables intention prediction and vision-conditioned generation, advancing controllable motion synthesis and intention understanding. The code is available at https://github.com/JunyuShi02/MoGIC
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。