arXiv:2605.15195cs.CV2026-05被引 7

VGGT-Ω大幅提升重建精度与效率,支持动态场景建模。

VGGT-$Ω$

论文配图:VGGT-$Ω$
图 1 · 摘自论文原文
  • 简化架构,用单个预测头和寄存器注意力替代复杂模块。
  • 训练时仅需原模型30%显存,可使用15倍更多标注数据。
  • 在Sintel上相机估计准确率提升77%,适配视觉语言动作系统。

近期的前馈重建模型(如VGGT)已展现出与传统优化方法相当的性能,并提供利于其他任务的几何感知特征。本文提出VGGT-Ω,显著提升静态与动态场景的重建精度、效率与能力。为实现前所未有的训练规模,我们引入架构改进以提高训练效率,构建支持动态场景的高质量数据标注流程,并设计自监督学习协议。通过采用单一密集预测头与多任务监督,移除高分辨率卷积层,利用寄存器聚合场景信息并引入寄存器注意力,限制帧间信息交换仅在寄存器间进行,部分替代全局注意力。训练中仅需原模型约30%的GPU内存,使我们能使用15倍于以往的监督数据,并充分利用大量未标注视频数据。VGGT-Ω在多个基准上取得优异重建结果,例如在Sintel上相机估计准确率提升77%。此外,所学寄存器还能增强视觉-语言-动作模型,支持与语言对齐,表明重建可作为空间理解的强大可扩展代理任务。

原文摘要 · Abstract (English)

Recent feed-forward reconstruction models, such as VGGT, have proven competitive with traditional optimization-based reconstructors while also providing geometry-aware features useful for other tasks. Here, we show that the quality of these models scales predictably with model and data size. We do so by introducing VGGT-$Ω$, which substantially improves reconstruction accuracy, efficiency, and capabilities for both static and dynamic scenes. To enable training this model at an unprecedented scale, we introduce architectural changes that improve training efficiency, a high-quality data annotation pipeline that supports dynamic scenes, and a self-supervised learning protocol. We simplify VGGT's architecture by using a single dense prediction head with multi-task supervision and removing the expensive high-resolution convolutional layers. We also use registers to aggregate scene information into a compact representation and introduce register attention, which restricts inter-frame information exchange to these registers, in part replacing global attention. In this way, during training, VGGT-$Ω$ uses only about 30% of the GPU memory of its predecessor, allowing us to train with 15x more supervised data than prior work and to leverage vast amounts of unlabeled video data. VGGT-$Ω$ achieves strong results for reconstruction of static and dynamic scenes across multiple benchmarks, for example, improving over the previous best camera estimation accuracy on Sintel by 77%. We also show that the learned registers can improve vision-language-action models and support alignment with language, suggesting that reconstruction can be a powerful and scalable proxy task for spatial understanding. Project Page: http://vggt-omega.github.io/

三维重建自监督学习动态场景模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。