arXiv:2606.11568cs.CV2026-06被引 1

构建4D动态感知问答数据集,提升视觉语言模型对运动的理解能力

4DP-QA: Scalable QA for 4D Perception in Vision Language Models

论文配图:4DP-QA: Scalable QA for 4D Perception in Vision Language Models
图 1 · 摘自论文原文
  • 提出真运动追踪技术,分离相机与物体运动,实现直观动态描述
  • 构建40万样本训练集和2200样本评测集,显著提升模型运动理解性能
  • 适合研究视频理解、动态场景建模及多模态推理的学者使用

尽管已有进展,视觉语言模型仍难以把握世界的动态特性。我们发现,理解4D场景本身已具挑战性,且受两大因素影响:一是模型通过2D图像投影间接观察运动;二是现有数据集未能区分物体与相机运动。为此,我们提出一种面向运动感知的问答生成流程,特别通过传统跟踪与新型固定参考系下的真运动追踪,有效解耦相机与物体运动,提供直观的运动描述。基于该流程,我们生成了一个包含40万样本的大规模训练数据集4DP-QA,以及一个2200样本的基准测试集4DP-QA-Bench。在外部基准上,使用该数据集训练现有模型可获得性能提升,验证了方法的有效性。

原文摘要 · Abstract (English)

Despite recent advances, Vision Language Models (VLMs) still struggle to grasp the dynamics of the world. We note that the ability to reason about a 4D scene, challenging in itself, is further complicated by two factors. First, VLMs observe motion indirectly via its projection onto 2D images. Second, existing datasets fail to disentangle object and camera motion. To address these challenges, we present a QA generation pipeline that focuses on motion-related scene understanding. We take particular care of the entanglement of camera and object motion by casting tracking in both the traditional way and in a novel, fixed reference system, dubbed True-Motion Tracking, which provides an intuitive description of motion. From this pipeline, we generate a large-scale training dataset of 400K samples, 4DP-QA (4D Perception QA), and a 2.2K-sample benchmark, 4DP-QA-Bench. Training existing models on our dataset yields performance improvements on an external benchmark, validating the effectiveness of our method.

4D感知视觉语言模型运动理解问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。