arXiv:2604.05689cs.CVcs.AI2026-04中稿 · CVPR被引 2

CRFT通过统一框架实现跨模态图像精准配准,尤其擅长处理大幅形变。

CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registration

论文配图:CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registration
图 1 · 摘自论文原文
  • 基于变换器架构学习跨模态不变特征流,联合对齐与光流估计
  • 多尺度相关与分层融合使配准在大尺度变化下仍保持精度
  • 适合医学影像、遥感、自动驾驶等需跨模态空间对应场景

我们提出一致递归特征流变换器(CRFT),一种基于特征流学习的统一粗到精框架,用于鲁棒的跨模态图像配准。CRFT在变换器架构中学习一种模态无关的特征流表示,联合完成特征对齐与光流估计。粗粒度阶段通过多尺度特征相关建立全局对应关系,细粒度阶段则通过分层特征融合与自适应空间推理细化局部细节。为增强几何适应性,采用迭代差异引导注意力机制结合空间几何变换(SGT),递归优化光流场,逐步捕捉细微空间不一致性并强制特征级一致性。该设计在大幅仿射与尺度变化下仍能实现精准对齐,并保持模态间结构连贯性。在多种跨模态数据集上的大量实验表明,CRFT在准确性和鲁棒性上持续优于现有最先进方法。除配准外,CRFT还提供了一种可泛化的多模态空间对应范式,适用于遥感、自主导航和医学成像等领域。代码与数据集公开于 https://github.com/NEU-Liuxuecong/CRFT。

原文摘要 · Abstract (English)

We present Consistent-Recurrent Feature Flow Transformer (CRFT), a unified coarse-to-fine framework based on feature flow learning for robust cross-modal image registration. CRFT learns a modality-independent feature flow representation within a transformer-based architecture that jointly performs feature alignment and flow estimation. The coarse stage establishes global correspondences through multi-scale feature correlation, while the fine stage refines local details via hierarchical feature fusion and adaptive spatial reasoning. To enhance geometric adaptability, an iterative discrepancy-guided attention mechanism with a Spatial Geometric Transform (SGT) recurrently refines the flow field, progressively capturing subtle spatial inconsistencies and enforcing feature-level consistency. This design enables accurate alignment under large affine and scale variations while maintaining structural coherence across modalities. Extensive experiments on diverse cross-modal datasets demonstrate that CRFT consistently outperforms state-of-the-art registration methods in both accuracy and robustness. Beyond registration, CRFT provides a generalizable paradigm for multimodal spatial correspondence, offering broad applicability to remote sensing, autonomous navigation, and medical imaging. Code and datasets are publicly available at https://github.com/NEU-Liuxuecong/CRFT.

图像配准跨模态变换器遥感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。