arXiv:2506.03148cs.CV2025-06CVPR被引 3

无需配准数据,自动对齐不同模态图像的像素点

Self-Supervised Spatial Correspondence Across Modalities

  • 用对比随机游走学习跨模态与同模态一致特征
  • 在无标注数据下实现RGB-深度、RGB-热成像等跨模态匹配
  • 适合多模态视觉任务,如遥感、机器人感知

我们提出一种方法,用于寻找不同视觉模态间的时空对应关系。给定来自不同模态的两张图像(如RGB图像和深度图),模型能识别出哪些像素对对应于场景中相同的物理点。为此,我们将对比随机游走框架扩展至同时学习跨模态与同模态的一致特征表示。所提模型结构简单,无需显式的光度一致性假设,可完全使用未标注数据训练,无需任何空间对齐的多模态图像对。我们在几何与语义对应任务上评估该方法:几何匹配涵盖挑战性任务如RGB-to-depth、RGB-to-thermal及其反向匹配;语义匹配则在照片-素描和跨风格图像对齐任务上测试。结果表明,该方法在所有基准上均表现优异。

原文摘要 · Abstract (English)

We present a method for finding cross-modal space-time correspondences. Given two images from different visual modalities, such as an RGB image and a depth map, our model identifies which pairs of pixels correspond to the same physical points in the scene. To solve this problem, we extend the contrastive random walk framework to simultaneously learn cycle-consistent feature representations for both cross-modal and intra-modal matching. The resulting model is simple and has no explicit photo-consistency assumptions. It can be trained entirely using unlabeled data, without the need for any spatially aligned multimodal image pairs. We evaluate our method on both geometric and semantic correspondence tasks. For geometric matching, we consider challenging tasks such as RGB-to-depth and RGB-to-thermal matching (and vice versa); for semantic matching, we evaluate on photo-sketch and cross-style image alignment. Our method achieves strong performance across all benchmarks.

跨模态对齐自监督学习多模态视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。