arXiv:2605.12027cs.CV2026-05

提出无需训练的动态静态解耦方法,提升单目视频4D重建精度。

4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation

论文配图:4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation
图 1 · 摘自论文原文
  • 通过分步解耦相机运动与物体运动,稳定姿态估计。
  • 在标准基准上点云指标显著提升,无需微调即具竞争力。
  • 适合关注4D重建与几何先验融合的研究者。

从单目视频重建动态4D场景是基础但极具挑战的任务。尽管近期3D基础模型提供了强大的几何先验,其在动态环境中的性能仍大幅下降,根源在于全局注意力机制中相机自运动与物体运动的固有耦合。本文提出一种无需训练的渐进式解耦框架,以粗到精的方式解离动态与静态成分。核心思路是先稳定相机位姿,再进行几何精修。具体包含三个协同模块:(1) 动态掩码引导的姿态解耦模块,隔离动态干扰,获得无运动参考帧;(2) 拓扑子空间手术机制,正交分解深度流形,在保留动态物体的同时向静态区域注入掩码感知的精细化几何;(3) 信息论置信度感知融合策略,将深度融合建模为异方差贝叶斯推断问题,通过逆方差加权自适应融合多轮预测。在标准4D重建基准上的大量实验表明,本方法在主要点云指标上实现一致且显著提升。值得注意的是,该方法无需微调即可展现竞争力,暗示数学严谨的动态-静态解耦具有潜在价值。

原文摘要 · Abstract (English)

Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance significantly degrades in dynamic environments. This degradation stems from a fundamental tension: the inherent coupling of camera ego-motion and object motion within global attention mechanisms. In this paper, we propose a novel, training-free progressive decoupling framework that disentangles dynamics from statics in a principled, coarse-to-fine manner. Our core insight is to resolve the tension by first stabilizing the camera pose, followed by geometric refinement. Specifically, our approach consists of three synergistic components: (1) a Dynamic-Mask-Guided Pose Decoupling module that isolates pose estimation from dynamic interference, yielding a stable motion-free reference frame; (2) a Topological Subspace Surgery mechanism that orthogonally decomposes the depth manifold, safely preserving dynamic objects while injecting refined, mask-aware geometry into static regions; and (3) an Information-Theoretic Confidence-Aware Fusion strategy that formulates depth integration as a heteroscedastic Bayesian inference problem, adaptively blending multi-pass predictions via inverse-variance weighting. Extensive experiments on standard 4D reconstruction benchmarks demonstrate that our method achieves consistent and substantial improvements across principal point-cloud metrics. Notably, our approach shows competitive performance in robust 4D scene reconstruction without requiring fine-tuning, suggesting the potential of mathematically grounded dynamic-static disentanglement.

4D重建视觉几何动态解耦Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。