让文字音乐生成舞蹈,实现精细动作控制。
Motion Anything: Any to Motion Generation
- 用注意力机制动态标记关键帧和身体部位,提升控制精度。
- 在HumanML3D上FID降低15%,多数据集表现超越现有方法。
- 支持文本+音乐多模态输入,适合舞蹈生成与创意设计者。
条件化动作生成在计算机视觉中研究广泛,但仍面临两大挑战:一是掩码自回归方法虽优于扩散模型,但现有掩码模型缺乏根据条件动态优先处理关键帧与身体部位的机制;二是不同条件模态(如文本、音乐)难以有效融合,影响生成动作的可控性与连贯性。为此,我们提出Motion Anything,一种多模态动作生成框架,引入基于注意力的掩码建模方法,实现对关键帧与动作的细粒度时空控制。模型可自适应编码文本、音乐等多模态条件,提升生成可控性。此外,我们构建了新数据集Text-Music-Dance(TMD),包含2,153对文本-音乐-舞蹈数据,规模为AIST++的两倍,填补了该领域的数据空白。大量实验表明,Motion Anything在多个基准上超越当前最优方法,在HumanML3D上实现15%的FID提升,并在AIST++和TMD上保持一致性能优势。
原文摘要 · Abstract (English)
Conditional motion generation has been extensively studied in computer vision, yet two critical challenges remain. First, while masked autoregressive methods have recently outperformed diffusion-based approaches, existing masking models lack a mechanism to prioritize dynamic frames and body parts based on given conditions. Second, existing methods for different conditioning modalities often fail to integrate multiple modalities effectively, limiting control and coherence in generated motion. To address these challenges, we propose Motion Anything, a multimodal motion generation framework that introduces an Attention-based Mask Modeling approach, enabling fine-grained spatial and temporal control over key frames and actions. Our model adaptively encodes multimodal conditions, including text and music, improving controllability. Additionally, we introduce Text-Music-Dance (TMD), a new motion dataset consisting of 2,153 pairs of text, music, and dance, making it twice the size of AIST++, thereby filling a critical gap in the community. Extensive experiments demonstrate that Motion Anything surpasses state-of-the-art methods across multiple benchmarks, achieving a 15% improvement in FID on HumanML3D and showing consistent performance gains on AIST++ and TMD. See our project website https://steve-zeyu-zhang.github.io/MotionAnything
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。