arXiv:2508.10838cs.CV2025-08被引 1

通过多基线几何一致性实现无监督立体匹配,提升遮挡区域的训练效果。

Unsupervised Stereo via Multi-Baseline Geometry-Consistent Self-Training

  • 用不同目标图像构建师生对,引入可见性差异增强遮挡区监督
  • 在KITTI 2015和2012上超越现有最优方法,取得新纪录
  • 适合自动驾驶立体视觉、无标注数据场景下的模型训练

光度损失与伪标签自训练是无监督训练立体网络的常用方法,但在遮挡区域均面临监督不准确的问题:前者缺乏有效对应点,后者伪标签不可靠。为此,我们提出S³框架,基于多基线几何一致性设计。不同于传统自训练中师生共享相同立体对,S³为师生分配不同目标图像,引入自然的可见性不对称性。学生视图中被遮挡的区域往往在教师视图中仍可见且可匹配,从而在光度监督失效区域也能生成可靠伪标签。教师的视差经基线重缩放后用于指导学生学习。进一步提出遮挡感知加权策略,降低教师遮挡区域不可靠监督的影响,并鼓励学生学习鲁棒的遮挡补全能力。为支持训练,我们使用CARLA模拟器构建了MBS20K多基线立体数据集。大量实验表明,S³在遮挡与非遮挡区域均提供有效监督,具有强泛化能力,在KITTI 2015和2012基准上超越现有最先进方法。

原文摘要 · Abstract (English)

Photometric loss and pseudo-label-based self-training are two widely used methods for training stereo networks on unlabeled data. However, they both struggle to provide accurate supervision in occluded regions. The former lacks valid correspondences, while the latter's pseudo labels are often unreliable. To overcome these limitations, we present S$^3$, a simple yet effective framework based on multi-baseline geometry consistency. Unlike conventional self-training where teacher and student share identical stereo pairs, S$^3$ assigns them different target images, introducing natural visibility asymmetry. Regions occluded in the student's view often remain visible and matchable to the teacher, enabling reliable pseudo labels even in regions where photometric supervision fails. The teacher's disparities are rescaled to align with the student's baseline and used to guide student learning. An occlusion-aware weighting strategy is further proposed to mitigate unreliable supervision in teacher-occluded regions and to encourage the student to learn robust occlusion completion. To support training, we construct MBS20K, a multi-baseline stereo dataset synthesized using the CARLA simulator. Extensive experiments demonstrate that S$^3$ provides effective supervision in both occluded and non-occluded regions, achieves strong generalization performance, and surpasses previous state-of-the-art methods on the KITTI 2015 and 2012 benchmarks.

无监督学习立体匹配遮挡处理自训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。