通过纹理模式图提升立体匹配精度与可解释性
Motif Channel Opened in a White-Box: Stereo Matching via Motif Correlation Graph
- 构建纹理模式相关图捕捉特征通道中的重复纹理
- 多频域特征融合恢复几何结构,中测排名第一
- 适合追求高精度与模型可解释性的视觉算法研究者
真实世界中的立体匹配应用(如自动驾驶)对安全性和准确性要求极高。然而,基于学习的方法在某些特征通道中会丢失几何结构,成为精确细节匹配的瓶颈,且因深度学习的黑箱特性缺乏可解释性。本文提出MoCha-V2,一种新型学习范式。MoCha-V2引入纹理模式相关图(MCG),捕捉特征通道中的重复纹理模式(motifs),重建几何结构,并以更可解释的方式学习。随后通过小波逆变换整合多频域特征,生成的纹理特征用于恢复立体匹配中的几何结构。实验表明,MoCha-V2在发布时位列中测基准第一名。代码已公开于https://github.com/ZYangChen/MoCha-Stereo。
原文摘要 · Abstract (English)
Real-world applications of stereo matching, such as autonomous driving, place stringent demands on both safety and accuracy. However, learning-based stereo matching methods inherently suffer from the loss of geometric structures in certain feature channels, creating a bottleneck in achieving precise detail matching. Additionally, these methods lack interpretability due to the black-box nature of deep learning. In this paper, we propose MoCha-V2, a novel learning-based paradigm for stereo matching. MoCha-V2 introduces the Motif Correlation Graph (MCG) to capture recurring textures, which are referred to as ``motifs" within feature channels. These motifs reconstruct geometric structures and are learned in a more interpretable way. Subsequently, we integrate features from multiple frequency domains through wavelet inverse transformation. The resulting motif features are utilized to restore geometric structures in the stereo matching process. Experimental results demonstrate the effectiveness of MoCha-V2. MoCha-V2 achieved 1st place on the Middlebury benchmark at the time of its release. Code is available at https://github.com/ZYangChen/MoCha-Stereo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。