无需训练的视觉几何变换器,显著提升立体视觉性能
StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision
- 基于冻结的VGGT模型,设计无训练特征调优流程
- 在KITTI基准上达到第一名,超越所有已发表方法
- 适合追求高精度立体匹配的科研与工业应用
随着三维设备的发展,立体视觉任务(如立体匹配和立体转换)成为关键研究方向。现有立体视觉主干网络通常依赖单目深度估计模型或通用预训练视觉模型,但这些模型大多未显式利用相机位姿监督。由于几何知识对立体视觉至关重要,缺乏显式空间约束导致现有架构性能受限。考虑到视觉几何基础变换器(VGGT)在大量3D先验(包括相机位姿)上预训练,我们探索其作为立体视觉骨干的潜力。然而实验发现,直接应用VGGT会导致特征提取中几何细节严重退化,与其双目视觉需求冲突。为此,我们提出StereoVGGT,一种专为立体视觉定制的特征主干。通过利用冻结的VGGT并引入无训练特征调整流程,有效缓解几何退化,并挖掘模型内嵌的相机标定知识。基于StereoVGGT的立体匹配网络在KITTI基准上取得第一,验证其作为立体视觉骨干的高度有效性。
原文摘要 · Abstract (English)
Driven by the advancement of 3D devices, stereo vision tasks including stereo matching and stereo conversion have emerged as a critical research frontier. Contemporary stereo vision backbones typically rely on either Monocular Depth Estimation models or general-purpose Pre-trained Vision Models. Crucially, these models are predominantly pretrained without explicit supervision of camera poses. Given that such geometric knowledge is indispensable for stereo vision, the absence of explicit spatial constraints constitutes a significant performance bottleneck for existing architectures. Recognizing that the Visual Geometry Grounded Transformer (VGGT) operates as a foundation model pre-trained on extensive 3D priors, including camera poses, we investigate its potential as a robust backbone for stereo vision tasks. Nevertheless, empirical results indicate that its direct application to stereo vision yields suboptimal performance. We observe that VGGT suffers from a more significant degradation of geometric details during feature extraction. Such characteristics conflict with the requirements of binocular stereo vision, thereby constraining its efficacy for relative tasks. To bridge this gap, we propose StereoVGGT, a feature backbone specifically tailored for stereo vision. By leveraging the frozen VGGT and introducing a training-free feature adjustment pipeline, we mitigate geometric degradation and harness the latent camera calibration knowledge embedded within the model. StereoVGGT-based stereo matching network achieved the $1^{st}$ rank among all published methods on the KITTI benchmark, validating that StereoVGGT serves as a highly effective backbone for stereo vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。