arXiv:2604.20393cs.CV2026-04

提升ViT立体匹配性能,兼顾全局与局部信息。

MLG-Stereo: ViT Based Stereo Matching with Multi-Stage Local-Global Enhancement

论文配图:MLG-Stereo: ViT Based Stereo Matching with Multi-Stage Local-Global Enhancement
图 1 · 摘自论文原文
  • 分阶段增强全局与局部特征,适应任意分辨率输入
  • 在Middlebury和KITTI-2015上表现优于现有方法
  • 适合需要高精度立体匹配的视觉系统开发者

随着深度学习的发展,基于视觉变换器(ViT)的立体匹配方法因其出色的鲁棒性和零样本能力取得了显著进展。然而,由于ViT在处理分辨率敏感性方面的局限性以及对局部信息关注不足,其在细节预测和任意分辨率图像处理方面仍弱于基于卷积神经网络(CNN)的方法。为此,我们提出MLG-Stereo,一个系统级的设计方案,将全局建模扩展至编码器之外的阶段。首先,设计多粒度特征网络,有效平衡全局上下文与局部几何信息,实现对任意分辨率图像的全面特征提取,并缩小训练与推理尺度间的差距。其次,构建局部-全局代价体积,捕捉局部相关与全局感知的匹配信息。最后,引入局部-全局引导循环单元,在全局信息指导下迭代优化局部视差。在多个基准数据集上的大量实验表明,MLG-Stereo在Middlebury和KITTI-2015基准上表现优异,超越同期领先方法,在KITTI-2012上也取得突出结果。

原文摘要 · Abstract (English)

With the development of deep learning, ViT-based stereo matching methods have made significant progress due to their remarkable robustness and zero-shot ability. However, due to the limitations of ViTs in handling resolution sensitivity and their relative neglect of local information, the ability of ViT-based methods to predict details and handle arbitrary-resolution images is still weaker than that of CNN-based methods. To address these shortcomings, we propose MLG-Stereo, a systematic pipeline-level design that extends global modeling beyond the encoder stage. First, we propose a Multi-Granularity Feature Network to effectively balance global context and local geometric information, enabling comprehensive feature extraction from images of arbitrary resolution and bridging the gap between training and inference scales. Then, a Local-Global Cost Volume is constructed to capture both locally-correlated and global-aware matching information. Finally, a Local-Global Guided Recurrent Unit is introduced to iteratively optimize the disparity locally under the guidance of global information. Extensive experiments are conducted on multiple benchmark datasets, demonstrating that our MLG-Stereo exhibits highly competitive performance on the Middlebury and KITTI-2015 benchmarks compared to contemporaneous leading methods, and achieves outstanding results in the KITTI-2012 dataset.

立体匹配ViT特征融合深度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。