arXiv:2512.05398cs.CV2025-12

用视觉语言模型识别动态物体,提升野外视频的3D结构理解。

The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos

论文配图:The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
图 1 · 摘自论文原文
  • 结合VLM与SAM2,无需特定训练即可识别动态物体。
  • 在真实和合成视频上实现顶尖运动分割性能。
  • 适合需要鲁棒3D重建的自动驾驶、VR等场景。

从野外视频中准确估计相机位姿、三维场景几何和物体运动,是传统运动结构方法长期面临的挑战,尤其在存在动态物体时。现有基于学习的方法通过训练运动估计算法来过滤动态物体并聚焦静态背景,但其性能受限于大规模运动分割数据集的缺乏,导致分割不准确,进而影响3D结构理解。本文提出动态先验(Dynamic Prior, ourmodel),通过利用视觉语言模型(VLM)的强大推理能力与SAM2的细粒度空间分割能力,无需任务特定训练即可鲁棒地识别动态物体。该方法可无缝集成至最先进的相机位姿优化、深度重建和4D轨迹估计流程中。在合成与真实视频上的大量实验表明, ourmodel不仅在运动分割任务上达到当前最优表现,还显著提升了3D结构理解的精度与鲁棒性。

原文摘要 · Abstract (English)

Estimating accurate camera poses, 3D scene geometry, and object motion from in-the-wild videos is a long-standing challenge for classical structure from motion pipelines due to the presence of dynamic objects. Recent learning-based methods attempt to overcome this challenge by training motion estimators to filter dynamic objects and focus on the static background. However, their performance is largely limited by the availability of large-scale motion segmentation datasets, resulting in inaccurate segmentation and, therefore, inferior structural 3D understanding. In this work, we introduce the Dynamic Prior (\ourmodel) to robustly identify dynamic objects without task-specific training, leveraging the powerful reasoning capabilities of Vision-Language Models (VLMs) and the fine-grained spatial segmentation capacity of SAM2. \ourmodel can be seamlessly integrated into state-of-the-art pipelines for camera pose optimization, depth reconstruction, and 4D trajectory estimation. Extensive experiments on both synthetic and real-world videos demonstrate that \ourmodel not only achieves state-of-the-art performance on motion segmentation, but also significantly improves accuracy and robustness for structural 3D understanding.

3D重建动态物体视觉语言模型运动分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。