arXiv:2511.19971cs.CV2025-11被引 23

无需训练,利用注意力机制挖掘动态线索,实现高效4D场景重建。

VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction

  • 从3D模型的全局注意力中挖掘时序动态线索,生成静态/动态分离掩码。
  • 在六大数据集上显著提升动态物体分割与相机位姿估计精度。
  • 支持超长序列单次推理,适合实时动态场景重建应用。

4D场景重建面临挑战,需有效分离动态物体与静态背景。尽管VGGT等3D基础模型能准确重建几何结构,但当动态物体占主导时性能显著下降。现有4D方法常依赖外部先验、复杂后优化或需在4D数据集上微调。本文提出VGGT4D,一种无需训练的框架,扩展VGGT以实现鲁棒的4D重建。核心发现是:VGGT的全局注意力层已隐式编码丰富的分层动态线索。我们通过格拉姆相似性挖掘并增强这些动态线索,并在时间窗口内聚合,生成静态与动态元素的分离掩码。为进一步锐化边界,引入基于投影梯度的精修策略。将精确掩码融入VGGT的早期推理阶段,有效缓解运动干扰对位姿估计与几何重建的影响。在六个数据集上,该方法在动态物体分割、相机位姿估计和稠密重建方面均表现优异,且支持超过500帧序列的单次推理。

原文摘要 · Abstract (English)

Reconstructing dynamic 4D scenes is challenging, as it requires robust disentanglement of dynamic objects from the static background. While 3D foundation models like VGGT provide accurate 3D geometry, their performance drops markedly when moving objects dominate. Existing 4D approaches often rely on external priors, heavy post-optimization, or require fine-tuning on 4D datasets. In this paper, we propose VGGT4D, a training-free framework that extends the 3D foundation model VGGT for robust 4D scene reconstruction. Our approach is motivated by the key finding that VGGT's global attention layers already implicitly encode rich, layer-wise dynamic cues. To obtain masks that decouple static and dynamic elements, we mine and amplify global dynamic cues via gram similarity and aggregate them across a temporal window. To further sharpen mask boundaries, we introduce a refinement strategy driven by projection gradient. We then integrate these precise masks into VGGT's early-stage inference, effectively mitigating motion interference in both pose estimation and geometric reconstruction. Across six datasets, our method achieves superior performance in dynamic object segmentation, camera pose estimation, and dense reconstruction. It also supports single-pass inference on sequences longer than 500 frames.

4D重建视觉几何动态分割自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。