arXiv:2602.08430cs.CV2026-02

改进注意力模型匹配局部特征,让模型通用且无需重新训练。

Understanding and Optimizing Attention-Based Sparse Matching for Diverse Local Features

  • 用多源关键点微调现有模型,提升跨检测器泛化能力。
  • 发现检测器性能差异比描述子影响更大,是决定匹配精度的关键。
  • 新模型零样本适配新检测器,效果不输专门训练的模型。

我们重新审视基于注意力机制的稀疏图像匹配模型在多种局部特征上的训练问题。首先识别出一个长期被忽视的关键设计选择,显著影响LightGlue模型性能。接着研究检测器与描述子在Transformer匹配框架中的作用,发现检测器而非描述子通常是性能差异的主要原因。最后提出一种新方法:利用多样检测器生成的关键点对现有匹配模型进行微调,构建出通用、检测器无关的模型。该模型在部署为新检测器的零样本匹配器时,达到或超过专为这些特征训练的模型精度。研究结果为Transformer匹配模型的应用和局部特征设计提供了重要启示。

原文摘要 · Abstract (English)

We revisit the problem of training attention-based sparse image matching models for various local features. We first identify one critical design choice that has been previously overlooked, which significantly impacts the performance of the LightGlue model. We then investigate the role of detectors and descriptors within the transformer-based matching framework, finding that detectors, rather than descriptors, are often the primary cause for performance difference. Finally, we propose a novel approach to fine-tune existing image matching models using keypoints from a diverse set of detectors, resulting in a universal, detector-agnostic model. When deployed as a zero-shot matcher for novel detectors, the resulting model achieves or exceeds the accuracy of models specifically trained for those features. Our findings offer valuable insights for the deployment of transformer-based matching models and the future design of local features.

图像匹配注意力机制通用模型关键点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。