arXiv:2511.05369cs.CV2025-11中稿 · 3DV 2026被引 2

为复杂人体动作序列生成带时间定位的详细描述。

Dense Motion Captioning

  • 设计新任务与数据集,实现动作的精准时序标注。
  • 构建含6万条复杂动作序列的CompMo数据集,每条包含2-10个动作。
  • 提出DEMO模型,结合大语言模型与运动适配器,效果显著领先。

3D人体动作与语言的融合研究多集中于文本到动作生成,而动作理解仍不充分。本文提出密集动作描述(Dense Motion Captioning)新任务,旨在对3D人体动作序列中的行为进行时序定位并生成描述。现有数据集缺乏精细时序标注,且多为短序列、动作少。为此,我们构建了首个大规模复杂动作数据集CompMo,通过精心设计的数据生成流程,包含60,000条动作序列,每条由至少2个至最多10个动作组成,均带有精确的时间边界标注。我们还提出了DEMO模型,将大语言模型与轻量级运动适配器结合,训练以生成密集且时序对齐的描述。实验表明,DEMO在CompMo及适配基准上均显著优于现有方法,为未来3D动作理解与描述研究奠定了坚实基线。

原文摘要 · Abstract (English)

Recent advances in 3D human motion and language integration have primarily focused on text-to-motion generation, leaving the task of motion understanding relatively unexplored. We introduce Dense Motion Captioning, a novel task that aims to temporally localize and caption actions within 3D human motion sequences. Current datasets fall short in providing detailed temporal annotations and predominantly consist of short sequences featuring few actions. To overcome these limitations, we present the Complex Motion Dataset (CompMo), the first large-scale dataset featuring richly annotated, complex motion sequences with precise temporal boundaries. Built through a carefully designed data generation pipeline, CompMo includes 60,000 motion sequences, each composed of multiple actions ranging from at least two to ten, accurately annotated with their temporal extents. We further present DEMO, a model that integrates a large language model with a simple motion adapter, trained to generate dense, temporally grounded captions. Our experiments show that DEMO substantially outperforms existing methods on CompMo as well as on adapted benchmarks, establishing a robust baseline for future research in 3D motion understanding and captioning.

动作理解时序标注语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。