arXiv:2512.10927cs.CV2025-12被引 5

用自动化方法生成细粒度运动数据,提升视频理解模型的物理推理能力。

FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos

  • 通过目标检测与轨迹追踪,结合大语言模型自动生成运动描述和问答对。
  • 在多个基准上超越Gemini-2.5 Flash等闭源模型,且不损害其他任务性能。
  • 适合需要提升空间运动推理能力的研究者,尤其关注视频理解与物理建模。

运动理解是物理推理的基础,使模型能够推断动态并预测未来状态。然而,当前先进模型在最新运动基准测试中仍表现不佳,主要因缺乏大规模、细粒度的运动数据集。现有数据集多依赖昂贵的人工标注,严重制约可扩展性。为此,我们提出FoundationMotion,一个全自动的数据构建流程,用于生成大规模运动数据集。该方法首先在视频中检测并跟踪物体以提取轨迹,再结合轨迹信息与视频帧,利用大语言模型(LLMs)生成细粒度描述及多样化的运动与空间推理问答对。基于此流程生成的数据集,我们微调了NVILA-Video-15B和Qwen2.5-7B等开源模型,在不牺牲其他任务性能的前提下,显著提升运动理解能力。值得注意的是,我们的模型在多个运动理解数据集和基准上超越了Gemini-2.5 Flash等强闭源基线,以及Qwen2.5-VL-72B等大型开源模型。FoundationMotion为构建细粒度运动数据集提供了可扩展方案,有效赋能多种模型的运动理解与空间推理能力。

原文摘要 · Abstract (English)

Motion understanding is fundamental to physical reasoning, enabling models to infer dynamics and predict future states. However, state-of-the-art models still struggle on recent motion benchmarks, primarily due to the scarcity of large-scale, fine-grained motion datasets. Existing motion datasets are often constructed from costly manual annotation, severely limiting scalability. To address this challenge, we introduce FoundationMotion, a fully automated data curation pipeline that constructs large-scale motion datasets. Our approach first detects and tracks objects in videos to extract their trajectories, then leverages these trajectories and video frames with Large Language Models (LLMs) to generate fine-grained captions and diverse question-answer pairs about motion and spatial reasoning. Using datasets produced by this pipeline, we fine-tune open-source models including NVILA-Video-15B and Qwen2.5-7B, achieving substantial improvements in motion understanding without compromising performance on other tasks. Notably, our models outperform strong closed-source baselines like Gemini-2.5 Flash and large open-source models such as Qwen2.5-VL-72B across diverse motion understanding datasets and benchmarks. FoundationMotion thus provides a scalable solution for curating fine-grained motion datasets that enable effective fine-tuning of diverse models to enhance motion understanding and spatial reasoning capabilities.

视频理解运动推理自动生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。