构建124万动作序列的视频运动基准,支持上下文动作生成与理解。
MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations
- 整合13个数据集,构建含1.24M动作序列的大规模运动库。
- 自动生成基于运动学规则的解耦文本描述,提升对齐精度。
- 适用于人体动作生成、上下文动作建模等任务,适合研究者使用。
本文针对大规模运动模型(LMM)的构建与评估难题,提出MotionBank——一个包含13个视频动作数据集、124万运动序列和1.329亿帧自然多样人体动作的大规模视频运动基准。与实验室采集的动作不同,真实场景视频蕴含丰富的上下文交互动作(如人与人、物、环境互动)。为实现更优的运动-文本对齐,我们设计了一种基于运动学特征的规则化、无偏且解耦的自动文本生成算法。大量实验表明,该数据集在人体动作生成、上下文动作生成与理解等任务上均具显著优势。视频动作与规则化文本注释可作为大模型训练的有效替代方案。数据集、代码与基准将公开于https://github.com/liangxuy/MotionBank。
原文摘要 · Abstract (English)
In this paper, we tackle the problem of how to build and benchmark a large motion model (LMM). The ultimate goal of LMM is to serve as a foundation model for versatile motion-related tasks, e.g., human motion generation, with interpretability and generalizability. Though advanced, recent LMM-related works are still limited by small-scale motion data and costly text descriptions. Besides, previous motion benchmarks primarily focus on pure body movements, neglecting the ubiquitous motions in context, i.e., humans interacting with humans, objects, and scenes. To address these limitations, we consolidate large-scale video action datasets as knowledge banks to build MotionBank, which comprises 13 video action datasets, 1.24M motion sequences, and 132.9M frames of natural and diverse human motions. Different from laboratory-captured motions, in-the-wild human-centric videos contain abundant motions in context. To facilitate better motion text alignment, we also meticulously devise a motion caption generation algorithm to automatically produce rule-based, unbiased, and disentangled text descriptions via the kinematic characteristics for each motion. Extensive experiments show that our MotionBank is beneficial for general motion-related tasks of human motion generation, motion in-context generation, and motion understanding. Video motions together with the rule-based text annotations could serve as an efficient alternative for larger LMMs. Our dataset, codes, and benchmark will be publicly available at https://github.com/liangxuy/MotionBank.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。