提升全景立体匹配的置信度与表面法向预测精度。
Revisiting Matching Response and Swept Feature Volumes for Wide-baseline Omnidirectional Stereo

- 用3D编码器解码器的响应期望值作为内在置信信号。
- 联合学习中引入扫掠特征重采样,提升深度一致性。
- 无需额外模块,适合自动驾驶等实际部署场景。
本文提出一种全景立体匹配中的置信度估计训练策略,针对宽基线设置下频繁出现的模糊匹配问题。通过重新解读3D编码器-解码器块生成的匹配响应,发现其期望值可作为内在置信信号。基于此,方法直接惩罚模糊响应,无需辅助头、多轮推理或额外模块,实现更高效且泛化性更强的预测。此外,引入扫掠特征体积重采样机制,将3D CNN生成的响应特征利用回归出的正匹配索引重采样后,由2D CNN预测如表面法向量等元信息。该联合学习过程引入辅助几何正则化,在响应聚合阶段利用额外上下文线索,提升深度一致性。实验表明,本方法在增强置信度估计和表面法向预测性能的同时,保持了面向自主移动应用的部署实用性。
原文摘要 · Abstract (English)
In this paper, we propose a training strategy for confidence estimation in omnidirectional stereo, targeting the ambiguous matches that frequently occur in wide-baseline setups. Reinterpreting the matching responses produced by the 3D encoder decoder block, we show that their expectation values provide intrinsic confidence signals. Building on this, our method directly penalizes ambiguous responses without auxiliary heads, multi-pass inference, or additional modules, resulting in more efficient and generalized predictions. Beyond confidence, we introduce swept feature volume resampling, where response features produced by 3D CNNs are resampled using regressed positive matching indices and then processed by 2D CNNs to predict meta-information such as surface normals. This joint learning introduces auxiliary geometric regularization and improves depth coherence by leveraging additional contextual cues during response aggregation stage. Experimental results demonstrate that our approach enhances both confidence estimation and surface normal prediction while maintaining deployment practicality for autonomous mobility applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。