融合多源特征提升两视图匹配准确率,解决重复结构下的误匹配问题。
See More, Match Better: Multi-Source Feature Fusion for Two-View Correspondence Learning

- 融合几何、纹理语义与结构语义三类特征进行匹配判断。
- 在MegaDepth和HPatches数据集上优于现有方法,尤其在复杂场景下表现更优。
- 适合需要高精度匹配的三维重建与视觉定位任务。
两视图对应关系学习旨在通过利用图像对间潜在差异,区分真实对应点(内点)与虚假对应点(外点)。现有方法主要依赖基于坐标的几何一致性,但在包含重复结构、无纹理区域或局部相似几何模式的场景中,常因伪一致性外点而失效。为此,本文提出TriMatch,一种用于两视图对应关系学习的多源特征融合框架,包含特征提取与特征精炼两部分。在特征提取阶段,TriMatch联合提取几何、纹理语义与结构语义特征,提供互补证据以支持对应关系判别。为弥合语义与几何特征间的差距,引入专用的纹理-几何对齐与结构-几何对齐模块,分别实现纹理与结构语义特征向几何特征的映射。进一步设计语义引导的对应关系调制模块,利用语义信息调制几何特征,抑制在几何上合理但语义不一致的对应关系。在特征精炼阶段,采用分层语义增强的对应关系精炼策略,逐步建模对应关系依赖性并重校多上下文特征响应,从而实现更可靠的内点/外点判别。大量实验表明,TriMatch在有效性、鲁棒性与泛化能力方面均表现出色。
原文摘要 · Abstract (English)
Two-view correspondence learning aims to distinguish true correspondences (inliers) from false ones (outliers) in image pairs by leveraging their underlying differences. Existing methods mainly rely on coordinate-based geometric consistency. However, they often struggle with pseudo-consistent outliers in scenes containing repetitive structures, textureless regions, or locally similar geometric patterns. To address this limitation, we propose TriMatch, a multi-source feature fusion framework for two-view correspondence learning, which consists of two parts: feature extraction and feature refinement. In feature extraction, TriMatch jointly extracts geometric, texture semantic, and structural semantic features to provide complementary evidence for correspondence discrimination. To bridge the gap between semantic and geometric features, texture and structural semantic features are aligned with geometric features through dedicated Texture-Geometric Alignment and Structural-Geometric Alignment modules, respectively. We further introduce a Semantic-Guided Correspondence Modulation module, which modulates geometric features using semantic information to suppress geometrically plausible but semantically inconsistent correspondences. In feature refinement, a Hierarchical Semantic-Enhanced Correspondence Refinement strategy progressively models correspondence dependencies and recalibrates multi-context feature responses, enabling more reliable inlier-outlier discrimination. Extensive experiments demonstrate the effectiveness, robustness, and generalization capability of TriMatch.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。