arXiv:2604.09366cs.CV2026-04被引 2

通过不确定性先验提升动态4D场景重建的精度与鲁棒性。

Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors

  • 用信息熵引导注意力聚合,分离动态运动与语义噪声
  • 通过邻域一致性约束消除结构异常点,降低13.43%误差
  • 无需微调或逐场景优化,适合实时动态重建应用

动态4D场景重建是重要但极具挑战的任务。尽管如VGGT等3D基础模型在静态场景表现优异,但在动态序列中因运动带来的几何模糊而性能下降。为此,本文提出一种框架,通过在重建过程中建模不确定性来解耦动态与静态成分。引入三种协同机制:(1) 信息熵引导的子空间投影,利用信息论权重自适应聚合多头注意力分布,有效分离动态运动信号与语义噪声;(2) 局部一致性驱动的几何净化,通过半径邻域约束强化空间连续性,剔除结构异常点;(3) 不确定性感知的跨视角一致性,将多视角投影优化建模为异方差最大似然估计问题,以深度置信度作为概率权重。在动态基准测试上,本方法相比现有最优方法,平均精度误差降低13.43%,分割F-measure提升10.49%。框架保持前馈推理效率,无需任务特定微调或每场景优化。

原文摘要 · Abstract (English)

Reconstructing dynamic 4D scenes is an important yet challenging task. While 3D foundation models like VGGT excel in static settings, they often struggle with dynamic sequences where motion causes significant geometric ambiguity. To address this, we present a framework designed to disentangle dynamic and static components by modeling uncertainty across different stages of the reconstruction process. Our approach introduces three synergistic mechanisms: (1) Entropy-Guided Subspace Projection, which leverages information-theoretic weighting to adaptively aggregate multi-head attention distributions, effectively isolating dynamic motion cues from semantic noise; (2) Local-Consistency Driven Geometry Purification, which enforces spatial continuity via radius-based neighborhood constraints to eliminate structural outliers; and (3) Uncertainty-Aware Cross-View Consistency, which formulates multi-view projection refinement as a heteroscedastic maximum likelihood estimation problem, utilizing depth confidence as a probabilistic weight. Experiments on dynamic benchmarks show that our approach outperforms current state-of-the-art methods, reducing Mean Accuracy error by 13.43\% and improving segmentation F-measure by 10.49\%. Our framework maintains the efficiency of feed-forward inference and requires no task-specific fine-tuning or per-scene optimization.

4D重建不确定性建模动态场景视觉几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。