arXiv:2411.07430cs.CV2024-11被引 5

XPoint无需标注数据,可自适应匹配多种波段图像。

XPoint: A Self-Supervised Visual-State-Space based Architecture for Multispectral Image Registration

  • 用自监督方法生成抗视角和波段变化的伪关键点
  • 在5个数据集上均超越或持平当前最佳结果
  • 适合需要快速适配新波段的图像配准场景

多光谱图像匹配因波段间非线性亮度差异、极端视角变化及标注数据稀缺而面临挑战。现有方法通常仅针对特定波段(如可见-红外)设计,依赖深度图或相机位姿等昂贵监督信号,难以跨模态迁移。为此,我们提出XPoint,一种自监督、模块化的图像匹配框架,支持在对齐的多光谱数据集上快速定制与微调。该框架采用可配置组件,通过基检测器生成对视角和波段变化不变的伪真值关键点。利用在分割任务上预训练的VMamba编码器提取鲁棒特征,并配备三个联合解码头:两个用于兴趣点与描述子提取,一个任务特定的单应性回归头施加几何约束,显著提升图像配准性能。该架构可快速适配多种模态,已在光学-热成像数据上训练,并成功微调至视觉-近红外、视觉-红外、视觉-远红外及视觉-合成孔径雷达等场景。实验表明,XPoint在五个不同多光谱数据集上的特征匹配与图像配准任务中持续优于或匹配现有最佳方法。

原文摘要 · Abstract (English)

Accurate multispectral image matching presents significant challenges due to non-linear intensity variations across spectral modalities, extreme viewpoint changes, and the scarcity of labeled datasets. Current state-of-the-art methods are typically specialized for a single spectral difference, such as visibleinfrared, and struggle to adapt to other modalities due to their reliance on expensive supervision, such as depth maps or camera poses. To address the need for rapid adaptation across modalities, we introduce XPoint, a self-supervised, modular image-matching framework designed for adaptive training and fine-tuning on aligned multispectral datasets, allowing users to customize key components based on their specific tasks. XPoint employs modularity and self-supervision to allow for the adjustment of elements such as the base detector, which generates pseudoground truth keypoints invariant to viewpoint and spectrum variations. The framework integrates a VMamba encoder, pretrained on segmentation tasks, for robust feature extraction, and includes three joint decoder heads: two are dedicated to interest point and descriptor extraction; and a task-specific homography regression head imposes geometric constraints for superior performance in tasks like image registration. This flexible architecture enables quick adaptation to a wide range of modalities, demonstrated by training on Optical-Thermal data and fine-tuning on settings such as visual-near infrared, visual-infrared, visual-longwave infrared, and visual-synthetic aperture radar. Experimental results show that XPoint consistently outperforms or matches state-ofthe-art methods in feature matching and image registration tasks across five distinct multispectral datasets. Our source code is available at https://github.com/canyagmur/XPoint.

图像配准自监督学习多光谱模块化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。