arXiv:2502.05432cs.CVcs.LG2025-02被引 4

构建可理解复杂动作的通用动作模型,支持多种下游任务。

MoFM: A Large-Scale Human Motion Foundation Model

  • 用热立方体编码动作,将运动转化为离散单元表示。
  • 基于大规模动作数据训练,支持零样本、无监督等多任务场景。
  • 适合动作识别、生成与分析等各类动作相关应用。

基础模型(FM)因其可扩展性和跨任务泛化能力受到越来越多关注。受大语言模型进展启发,本文提出一种新型运动基础模型 MoFM,旨在实现对时空复杂人类动作的语义理解。为支持大规模训练,设计了 MotionBook——一个包含离散化动作的完整人体动作词典,利用热立方体捕捉时空运动热图,并借鉴离散变分模型原理,将人体运动编码为离散单元,实现更高效、可扩展的表示。MoFM 在大规模动作数据上训练,可作为多样化下游任务的基础骨干,支持零样本、无监督及有监督等多种范式。其灵活性使其适用于广泛的基于动作的应用场景。

原文摘要 · Abstract (English)

Foundation Models (FM) have increasingly drawn the attention of researchers due to their scalability and generalization across diverse tasks. Inspired by the success of FMs and the principles that have driven advancements in Large Language Models (LLMs), we introduce MoFM as a novel Motion Foundation Model. MoFM is designed for the semantic understanding of complex human motions in both time and space. To facilitate large-scale training, MotionBook, a comprehensive human motion dictionary of discretized motions is designed and employed. MotionBook utilizes Thermal Cubes to capture spatio-temporal motion heatmaps, applying principles from discrete variational models to encode human movements into discrete units for a more efficient and scalable representation. MoFM, trained on a large corpus of motion data, provides a foundational backbone adaptable to diverse downstream tasks, supporting paradigms such as one-shot, unsupervised, and supervised tasks. This versatility makes MoFM well-suited for a wide range of motion-based applications.

动作建模基础模型时空学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。