arXiv:2511.18373cs.CV2025-11被引 3

让视觉语言模型理解物体运动与空间关系,提升物理推理能力

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models

  • 通过深度编码和运动追踪,将物理动态转化为可理解的语义信号
  • 在4350段真实与AI生成视频上,实现接近闭源顶尖模型的物理推理性能
  • 适合需要精准物理理解的视频分析、智能教育等场景

视觉语言模型(VLMs)在标准视频任务上表现良好,但在涉及运动动力学和空间交互的物理推理任务中表现不佳。本文提出MASS——一种模型无关的方法,通过基于深度的3D编码和视觉定位,将时空信号注入VLM的语言空间,并结合运动追踪器捕捉物体动态。同时构建了MASS-Bench基准,包含4,350段真实与AIGC视频、8,361个自由形式的视频问答对,涵盖子片段级视觉检测、定位及全序列3D运动追踪标注。为增强跨模态对齐与推理能力,采用强化学习微调。实验表明,优化后的VLM在多个基线和更大模型上均表现更优,性能接近闭源最先进模型,仅比Gemini-2.5-Flash低2%。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) perform well on standard video tasks but struggle with physics-related reasoning involving motion dynamics and spatial interactions. We present a novel approach to address this gap by translating physical-world context cues into interpretable representations aligned with VLM perception, comprehension, and reasoning. We introduce MASS, a model-agnostic approach that injects spatiotemporal signals into the VLM language space via depth-based 3D encoding and visual grounding, coupled with a motion tracker for object dynamics. We also contribute a comprehensive benchmark, MASS-Bench, consisting of 4,350 real-world and AIGC videos and 8,361 free-form video question-answering pairs focused on physics-related comprehension tasks, with detailed annotations including visual detections and grounding over sub-segments, as well as full-sequence 3D motion tracking of entities. To strengthen cross-modal alignment and reasoning, we apply reinforcement fine-tuning to MASS. Experiments and ablations show that our refined VLMs outperform comparable baselines, larger models, and prior state-of-the-art models, achieving performance comparable to closed-source state-of-the-art VLMs, with only a 2\% gap to Gemini-2.5-Flash on physics reasoning and comprehension.

视觉语言模型物理推理时空建模运动追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。