arXiv:2512.20557cs.CV2025-12被引 6

让视觉语言模型学会动态空间推理,提升对物体运动与关系的理解能力。

Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models

  • 构建自动化数据生成管道,从真实视频中提取4D几何信息用于训练。
  • 引入轻量级几何选择模块,精准注入3D空间知识,显著提升推理性能。
  • 适配多场景复杂交互,适合需要动态空间理解的视觉模型研究者。

视觉语言模型(VLM)在通用理解方面表现优异,但在动态空间推理(DSR)上仍较薄弱,即难以分析物体在3D空间随时间演化的几何与关系,主要因缺乏可扩展的4D感知训练资源。为此,我们提出DSR Suite:首先设计自动化流水线,从真实视频中生成多选题问答对,利用现代视觉基础模型提取相机位姿、局部点云、物体掩码、朝向及3D轨迹等丰富几何与运动信息;这些信息构建了用于学习的DSR-Train和人工精修的评估基准DSR-Bench。相比以往工作,本数据强调(i)真实场景视频源,(ii)物体与场景级3D需求,(iii)视角变换,(iv)多物体交互,以及(v)细粒度、程序化答案。此外,提出轻量级几何选择模块(GSM),将问题语义与预训练4D重建先验中的相关知识压缩为紧凑的几何标记,避免冗余信息干扰。实验表明,将DSR-Train与GSM集成至Qwen2.5-VL-7B后,其动态空间推理能力显著提升,同时保持在通用视频理解基准上的准确率。

原文摘要 · Abstract (English)

Vision-language models (VLM) excel at general understanding yet remain weak at dynamic spatial reasoning (DSR), i.e., reasoning about the evolvement of object geometry and relationship in 3D space over time, largely due to the scarcity of scalable 4D-aware training resources. To bridge this gap across aspects of dataset, benchmark and model, we introduce DSR Suite. First, we propose an automated pipeline that generates multiple-choice question-answer pairs from in-the-wild videos for DSR. By leveraging modern vision foundation models, the pipeline extracts rich geometric and motion information, including camera poses, local point clouds, object masks, orientations, and 3D trajectories. These geometric cues enable the construction of DSR-Train for learning and further human-refined DSR-Bench for evaluation. Compared with previous works, our data emphasize (i) in-the-wild video sources, (ii) object- and scene-level 3D requirements, (iii) viewpoint transformations, (iv) multi-object interactions, and (v) fine-grained, procedural answers. Beyond data, we propose a lightweight Geometry Selection Module (GSM) to seamlessly integrate geometric priors into VLMs, which condenses question semantics and extracts question-relevant knowledge from pretrained 4D reconstruction priors into a compact set of geometry tokens. This targeted extraction avoids overwhelming the model with irrelevant knowledge. Experiments show that integrating DSR-Train and GSM into Qwen2.5-VL-7B significantly enhances its dynamic spatial reasoning capability, while maintaining accuracy on general video understanding benchmarks.

动态空间推理视觉语言模型4D理解几何先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。