arXiv:2409.02545cs.CV2024-09被引 3

用自监督预训练提升Transformer在立体匹配中的表现

UniTT-Stereo: Unified Training of Transformer for Enhanced Stereo Matching

  • 融合自监督预训练与监督学习,统一训练框架
  • 在KITTI和ETH3D上达到当前最佳性能
  • 适合研究立体匹配与Transformer架构的学者

尽管基于Transformer的方法在其他视觉任务中日益普及,立体深度估计仍以卷积方法为主。这主要受限于真实世界立体匹配标注数据稀缺,制约了基于Transformer方法的性能提升。本文提出UniTT-Stereo,通过统一自监督预训练与监督式立体匹配框架,最大化基于Transformer的立体匹配架构潜力。具体而言,从局部性归纳偏置角度出发,探索掩码图像区域重建与对应点预测的联合有效性。为应对重建与预测双重挑战,设计了一种随训练动态调整掩码比例的新策略,并结合专用于立体匹配的损失函数。在ETH3D、KITTI 2012和KITTI 2015等多个基准上验证了其先进性能。最后,通过特征图频谱分析及注意力图的局部性归纳偏置分析,揭示了该方法的优势。

原文摘要 · Abstract (English)

Unlike other vision tasks where Transformer-based approaches are becoming increasingly common, stereo depth estimation is still dominated by convolution-based approaches. This is mainly due to the limited availability of real-world ground truth for stereo matching, which is a limiting factor in improving the performance of Transformer-based stereo approaches. In this paper, we propose UniTT-Stereo, a method to maximize the potential of Transformer-based stereo architectures by unifying self-supervised learning used for pre-training with stereo matching framework based on supervised learning. To be specific, we explore the effectiveness of reconstructing features of masked portions in an input image and at the same time predicting corresponding points in another image from the perspective of locality inductive bias, which is crucial in training models with limited training data. Moreover, to address these challenging tasks of reconstruction-and-prediction, we present a new strategy to vary a masking ratio when training the stereo model with stereo-tailored losses. State-of-the-art performance of UniTT-Stereo is validated on various benchmarks such as ETH3D, KITTI 2012, and KITTI 2015 datasets. Lastly, to investigate the advantages of the proposed approach, we provide a frequency analysis of feature maps and the analysis of locality inductive bias based on attention maps.

立体匹配Transformer自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。