聚焦视频中特定区域的运动描述,提升细粒度理解能力。
MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

- 构建区域感知的运动描述系统,基于时空掩码精准定位运动区域。
- 生成15.9万条高质量运动描述数据,有效减少幻觉现象。
- 适用于需要精细动作理解的场景,如机器人视觉与视频分析。
我们提出MotionAtlas,一个面向以运动为中心的视频的详细描述系统,包含(1)专用的人工标注基准数据集,(2)可扩展的高质量训练样本构建流程,以及(3)一系列强大的Video-MLLM模型。与传统全局运动描述数据集不同,本工作聚焦于区域感知的运动描述:给定视频和时空掩码,模型生成目标区域内运动的精确描述,从而缓解视觉杂乱与运动混淆问题,并支持可靠、可量化的评估。具体而言,我们首先构建MotionAtlas-Bench,一个涵盖2,073个多项选择题的综合性基准,对精选的高质量运动中心视频进行细致标注,用于评估对象的细粒度运动理解能力。其次,设计了一种严格且可扩展的数据构建流程,利用自举式精炼机制抑制细粒度幻觉,生成15.9万条高质量运动描述数据。第三,设计了定制化的训练数据组合策略,在多种基线Video-MLLM上实现持续显著性能提升,包括Molmo2和Qwen3-VL。例如,MotionAtlas-4B在通用运动基准上平均超越Qwen3-VL-4B 5.2个百分点。该基准、数据集与代码均已公开。
原文摘要 · Abstract (English)
We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs. Unlike conventional global motion captioning datasets, we focus on region-aware motion captioning: given a video and a spatiotemporal mask, the model generates precise descriptions of motion within the target region, thereby alleviating visual clutter and motion entanglement and enabling reliable, quantifiable evaluation. Concretely, we first build MotionAtlas-Bench, a comprehensive benchmark comprising 2,073 multiple-choice questions, meticulously annotated for a curated set of high-quality, motion-centric videos, to evaluate fine-grained motion understanding of the objects in question. Second, we design a rigorous and scalable data pipeline that leverages self-bootstrap refinement to suppress fine-grained hallucinations, yielding 159k high-quality motion captioning data. Third, we design a tailored training data composition strategy, which achieves consistent and substantial performance gains across diverse baseline Video-MLLMs, including Molmo2 and Qwen3-VL. For instance, MotionAtlas-4B surpasses Qwen3-VL-4B by an average of 5.2 percentage points across general motion benchmarks. The benchmark, dataset, and code have been released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。