让AI看视频时理解运动逻辑并画出对应区域,提升对动态场景的精准感知。
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
- 通过问题引导模型进行隐式时空推理,生成像素级视频分割掩码。
- 在1715段视频上构建4类问题数据集,平均比现有模型高21.5%性能。
- 适合研究视频理解、运动推理与视觉定位的学者和开发者。
本文提出运动感知视频推理新任务,要求根据问题生成视觉答案(视频分割掩码),需具备隐式时空推理与定位能力。该任务拓展了传统显式动作/运动定位,通过问题实现更通用的隐式推理。为此,我们构建大规模数据集GROUNDMORE,包含1,715个视频片段、249,000个精心设计的对象掩码,涵盖因果、顺序、反事实和描述四类问题,用于评测深度运动推理能力。该数据集要求模型输出视觉答案,提供比文本更直观可解释的响应。同时引入新基线模型MORA,融合多模态大模型的推理能力、SAM的像素级感知能力及轻量级定位头的时序感知能力,在GROUNDMORE上表现优于现有最佳视觉定位模型平均21.5%。本工作旨在推动基于视频推理分割的鲁棒通用运动理解发展。
原文摘要 · Abstract (English)
In this paper, we introduce Motion-Grounded Video Reasoning, a new motion understanding task that requires generating visual answers (video segmentation masks) according to the input question, and hence needs implicit spatiotemporal reasoning and grounding. This task extends existing spatiotemporal grounding work focusing on explicit action/motion grounding, to a more general format by enabling implicit reasoning via questions. To facilitate the development of the new task, we collect a large-scale dataset called GROUNDMORE, which comprises 1,715 video clips, 249K object masks that are deliberately designed with 4 question types (Causal, Sequential, Counterfactual, and Descriptive) for benchmarking deep and comprehensive motion reasoning abilities. GROUNDMORE uniquely requires models to generate visual answers, providing a more concrete and visually interpretable response than plain texts. It evaluates models on both spatiotemporal grounding and reasoning, fostering to address complex challenges in motion-related video reasoning, temporal perception, and pixel-level understanding. Furthermore, we introduce a novel baseline model named Motion-Grounded Video Reasoning Assistant (MORA). MORA incorporates the multimodal reasoning ability from the Multimodal LLM, the pixel-level perception capability from the grounding model (SAM), and the temporal perception ability from a lightweight localization head. MORA achieves respectable performance on GROUNDMORE outperforming the best existing visual grounding baseline model by an average of 21.5% relatively. We hope this novel and challenging task will pave the way for future advancements in robust and general motion understanding via video reasoning segmentation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。