用高质量长文本生成更自然的人体动作,支持复杂表达。
SnapMoGen: Human Motion Generation from Expressive Texts
- 构建新数据集SnapMoGen,含44小时连续动作与详细长文本描述
- 模型MoMask++在两项基准上达最新水平,支持长序列动作生成
- 结合大模型重写用户输入,适配高表达性文本风格
近年来,文本到动作生成取得了显著进展。然而,现有方法仍受限于短句或通用文本提示,主要受制于数据集规模。这一限制削弱了细粒度控制能力,并影响对未见提示的泛化性能。本文提出SnapMoGen,一个新型文本-动作数据集,包含高质量动作捕捉数据与精确、富有表现力的文本标注。该数据集包含20,000段动作片段,总计44小时,配有122,000条详细文本描述,平均每条48词(远高于HumanML3D的12词)。重要的是,这些动作片段保留原始时间连续性,源自长序列数据,有利于长期动作生成与融合研究。我们还改进了以往的生成式掩码建模方法。提出的MoMask++将动作转换为多尺度标记序列,更充分挖掘标记容量,并使用单一生成式掩码Transformer学习生成所有标记。MoMask++在HumanML3D和SnapMoGen两个基准上均达到当前最优性能。此外,我们通过引入大语言模型重写用户输入,使其符合SnapMoGen的表达风格与叙述方式,实现对日常提示的有效处理。
原文摘要 · Abstract (English)
Text-to-motion generation has experienced remarkable progress in recent years. However, current approaches remain limited to synthesizing motion from short or general text prompts, primarily due to dataset constraints. This limitation undermines fine-grained controllability and generalization to unseen prompts. In this paper, we introduce SnapMoGen, a new text-motion dataset featuring high-quality motion capture data paired with accurate, expressive textual annotations. The dataset comprises 20K motion clips totaling 44 hours, accompanied by 122K detailed textual descriptions averaging 48 words per description (vs. 12 words of HumanML3D). Importantly, these motion clips preserve original temporal continuity as they were in long sequences, facilitating research in long-term motion generation and blending. We also improve upon previous generative masked modeling approaches. Our model, MoMask++, transforms motion into multi-scale token sequences that better exploit the token capacity, and learns to generate all tokens using a single generative masked transformer. MoMask++ achieves state-of-the-art performance on both HumanML3D and SnapMoGen benchmarks. Additionally, we demonstrate the ability to process casual user prompts by employing an LLM to reformat inputs to align with the expressivity and narration style of SnapMoGen. Project webpage: https://snap-research.github.io/SnapMoGen/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。