arXiv:2410.03311cs.CVcs.LG2024-10ICML被引 20

构建百万级动作数据集,训练出能生成多样动作的大型模型。

Scaling Large Motion Models with Million-Level Human Motions

  • 构建百万级动作数据集MotionLib,含分层文本描述。
  • 模型在未见过的动作上表现稳健,验证数据与模型规模协同提升效果。
  • 提出Motionbook编码方法,高效保留动作细节并增强表征能力。

受大语言模型成功的启发,人体动作理解领域正转向开发大型动作模型。尽管已有进展,但当前工作仍远未达到真正通用模型的标准,主要受限于高质量大规模数据的缺乏。为此,我们提出MotionLib,首个百万级动作生成数据集,规模至少是现有数据集的15倍,并包含分层文本描述。基于MotionLib,我们训练了大型动作模型 extit{projname},在多种人体活动中表现出稳健性能,包括未见动作。通过系统性研究,我们首次强调了数据与模型规模同步扩展对推进动作生成的重要性,并提供关键实现洞见。为更好融合动作模态,我们提出Motionbook,一种创新的动作编码方法,包含:(1) 紧凑且无损的动作特征表示;(2) 无2D查找的新型动作分词器,在保持精细动作细节的同时显著扩展码本容量,大幅增强动作令牌的表征能力。我们认为该工作为未来开发更通用、更强的运动生成模型奠定了基础。更多详情请访问 https://beingbeyond.github.io/Being-M0/。

原文摘要 · Abstract (English)

Inspired by the recent success of LLMs, the field of human motion understanding has increasingly shifted toward developing large motion models. Despite some progress, current efforts remain far from achieving truly generalist models, primarily due to the lack of massive high-quality data. To address this gap, we present MotionLib, the first million-level dataset for motion generation, which is at least 15$\times$ larger than existing counterparts and enriched with hierarchical text descriptions. Using MotionLib, we train a large motion model named \projname, demonstrating robust performance across a wide range of human activities, including unseen ones. Through systematic investigation, for the first time, we highlight the importance of scaling both data and model size for advancing motion generation, along with key insights to achieve this goal. To better integrate the motion modality, we propose Motionbook, an innovative motion encoding approach including (1) a compact yet lossless feature to represent motions; (2) a novel 2D lookup-free motion tokenizer that preserves fine-grained motion details while expanding codebook capacity, significantly enhancing the representational power of motion tokens. We believe this work lays the groundwork for developing more versatile and powerful motion generation models in the future. For further details, visit https://beingbeyond.github.io/Being-M0/.

动作生成大规模数据模型扩展编码方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。