arXiv:2606.16479cs.CVcs.AI2026-06中稿 · publication in the…被引 1

分析VGGT在DTU数据集上的不确定性质量,发现优化置信度阈值可提升重建精度。

Uncertainty Quality of VGGT: An Analysis on the DTU Benchmark Dataset

论文配图:Uncertainty Quality of VGGT: An Analysis on the DTU Benchmark Dataset
图 1 · 摘自论文原文
  • 通过设定有效置信度阈值过滤原始输出,提升可靠性。
  • 不确定性质量与3D重建精度正相关,优化后精度显著提高。
  • 适合关注3D重建可信度与鲁棒性评估的研究者。

视觉几何基础变换器(VGGT)在短时间内引发广泛关注,尤其因其荣获CVPR-2025最佳论文奖。与DUSt3R和MASt3R类似,VGGT旨在通过一个统一的前馈神经网络,直接从多张场景图像中快速预测相机位姿、深度图和稠密3D结构,从而取代传统的捆绑调整与特征匹配方法,实现秒级重建。其核心优势在于能以单次前向传播处理任意数量视角,无需后处理或迭代优化。这对摄影测量学而言,为实时、可扩展且易获取的3D重建开辟了新可能。在此背景下,不仅重建精度重要,高质量的不确定性估计同样关键,可增强模型可信度并支持稳健的质量保障。本文因此系统分析VGGT在DTU基准数据集上的不确定性预测质量,识别出有效的置信度阈值以过滤原始输出,并证明提升不确定性质量具有显著潜力,能进一步改善3D重建精度。

原文摘要 · Abstract (English)

Visual Geometry Grounded Transformer (VGGT) has already attracted a great deal of attention in a short period of time, not least due to the Best Paper Award at CVPR-2025. Similar to DUSt3R and MASt3R, VGGT aims to bring about a paradigm shift by replacing established methods like bundle adjustment and feature matching with a simple, unified, feed-forward neural network that predicts camera poses, depth maps, and dense 3D structure directly from multiple images of a scene in a few seconds. A key aspect is its ability to process an arbitrary number of views consistently in a single forward pass without any post-processing or iterative optimization. For photogrammetry, this opens new possibilities for real-time, scalable, and accessible 3D reconstruction. In this context, not only high reconstruction accuracy but also high-quality uncertainty estimates are crucial, as they foster trust and enable robust quality assurance. This paper therefore investigates the quality of VGGT's uncertainty predictions. The analysis identifies an effective confidence threshold for filtering VGGT's raw output and demonstrates that enhancing uncertainty quality holds strong potential for improving the accuracy of its 3D reconstructions.

3D重建不确定性估计神经网络视觉几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。