提出动态可见性感知的双视图亚像素级匹配方法,提升精度与效率。
CoMatch: Dynamic Covisibility-Aware Transformer for Bilateral Subpixel-Level Semi-Dense Image Matching
- 通过动态可见性评分自适应压缩特征令牌,减少冗余计算。
- 引入可见性辅助注意力机制,抑制非可见区域干扰,增强特征区分度。
- 双视图同时优化至亚像素级,适用于关键点敏感任务。
本文提出CoMatch,一种具备动态可见性感知与双边亚像素精度的新型半密集图像匹配方法。首先,针对全局粗粒度特征图中相邻令牌表征相似导致计算冗余的问题,设计基于可见性评分的令牌压缩模块,动态估计可见性并自适应聚合令牌,兼顾计算效率与表征能力。其次,为避免大量不可见区域特征交互带来的干扰,引入可见性辅助注意力机制,选择性抑制非可见区域的无关信息传播,实现对相关区域的鲁棒紧凑注意力。第三,发现现有方法仅将目标视图关键点调整至亚像素级,源视图仍停留在粗粒度,信息不足,影响关键点定位敏感任务。为此,提出简单而有效的精细相关性模块,同时将源与目标视图匹配候选精炼至亚像素级别,显著提升性能。在多个公开基准上的充分实验验证了CoMatch在准确性、效率与泛化性方面的优越表现。
原文摘要 · Abstract (English)
This prospective study proposes CoMatch, a novel semi-dense image matcher with dynamic covisibility awareness and bilateral subpixel accuracy. Firstly, observing that modeling context interaction over the entire coarse feature map elicits highly redundant computation due to the neighboring representation similarity of tokens, a covisibility-guided token condenser is introduced to adaptively aggregate tokens in light of their covisibility scores that are dynamically estimated, thereby ensuring computational efficiency while improving the representational capacity of aggregated tokens simultaneously. Secondly, considering that feature interaction with massive non-covisible areas is distracting, which may degrade feature distinctiveness, a covisibility-assisted attention mechanism is deployed to selectively suppress irrelevant message broadcast from non-covisible reduced tokens, resulting in robust and compact attention to relevant rather than all ones. Thirdly, we find that at the fine-level stage, current methods adjust only the target view's keypoints to subpixel level, while those in the source view remain restricted at the coarse level and thus not informative enough, detrimental to keypoint location-sensitive usages. A simple yet potent fine correlation module is developed to refine the matching candidates in both source and target views to subpixel level, attaining attractive performance improvement. Thorough experimentation across an array of public benchmarks affirms CoMatch's promising accuracy, efficiency, and generalizability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。