arXiv:2509.24318cs.CV2025-09ICCV被引 2

用轻量状态空间模型高效建模图像语义对应关系

Similarity-Aware Selective State-Space Modeling for Semantic Correspondence

  • 基于相似性感知的选通扫描机制,降低计算开销
  • 在标准数据集上达到当前最优性能,保持高分辨率特征
  • 适合需要快速准确匹配图像语义的场景

建立图像间的语义对应是计算机视觉中的基础但极具挑战的任务。传统特征度量方法虽能增强视觉特征,却可能忽略复杂的相互关联关系;而近期的相关度量方法受限于处理4维相关图带来的高计算成本。我们提出MambaMatcher,一种新方法,通过选择性状态空间模型(SSMs)高效建模高维相关性。该方法采用源自Mamba的线性复杂度算法改进的相似性感知选通扫描机制,有效优化4维相关图,同时不牺牲特征图分辨率或感受野。在标准语义对应基准测试中,MambaMatcher实现了最先进性能。

原文摘要 · Abstract (English)

Establishing semantic correspondences between images is a fundamental yet challenging task in computer vision. Traditional feature-metric methods enhance visual features but may miss complex inter-correlation relationships, while recent correlation-metric approaches are hindered by high computational costs due to processing 4D correlation maps. We introduce MambaMatcher, a novel method that overcomes these limitations by efficiently modeling high-dimensional correlations using selective state-space models (SSMs). By implementing a similarity-aware selective scan mechanism adapted from Mamba's linear-complexity algorithm, MambaMatcher refines the 4D correlation map effectively without compromising feature map resolution or receptive field. Experiments on standard semantic correspondence benchmarks demonstrate that MambaMatcher achieves state-of-the-art performance.

语义对应状态空间模型图像匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。