arXiv:2605.12064cs.CV2026-05

用文本语义提升光学与雷达图像配准精度,尤其在大幅形变下表现更优。

TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images

论文配图:TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images
图 1 · 摘自论文原文
  • 引入遥感场景文本描述,通过视觉-文本交互增强跨模态特征
  • 多尺度特征提取+粗到精匹配流程,在大形变下仍保持高精度
  • 适合遥感图像配准任务,对复杂地形和异质数据有强适应性

现有基于深度学习的方法虽能捕捉光学与合成孔径雷达(SAR)图像的共享特征以实现空间对齐,但在大几何形变下仍面临挑战,因模型需同时处理跨模态外观差异与复杂空间变换。为此,本文提出一种文本语义辅助的跨模态图像配准框架TAR。TAR利用遥感场景与地表覆盖类别的文本语义先验,缓解模态差距并增强跨模态特征学习。其包含三个模块:多尺度视觉特征学习(MSFL)模块、文本辅助特征增强(TAFE)模块和粗到精密集匹配(CFDM)模块。MSFL从光学与SAR图像中提取多尺度视觉特征;TAFE构建与遥感场景及地表覆盖对象相关的文本描述,并通过冻结的RemoteCLIP文本编码器提取文本特征,经由视觉-文本交互机制增强高层视觉特征,提升粗匹配可靠性;CFDM则基于增强后的高层特征建立粗对应关系,并利用低层特征细化匹配位置。在跨模态遥感图像上的实验表明,TAR性能优于多个先进方法,在大几何形变下仍取得显著提升。

原文摘要 · Abstract (English)

Existing deep learning-based methods can capture shared features from optical and synthetic aperture radar (SAR) images for spatial alignment. However, optical-SAR registration remains challenging under large geometric deformations, because the model needs to simultaneously handle cross-modal appearance discrepancies and complex spatial transformations. To address this issue, this paper proposes a text semantic-assisted cross-modal image registration framework, named TAR, for optical and SAR images. TAR exploits text semantic priors from remote sensing scenes and land-cover categories to alleviate the modality gap and enhance cross-modal feature learning. TAR consists of three components: a multi-scale visual feature learning (MSFL) module, a text-assisted feature enhancement (TAFE) module, and a coarse-to-fine dense matching (CFDM) module. MSFL extracts multi-scale visual features from optical and SAR images. TAFE constructs text descriptors related to remote sensing scenes and land-cover objects, and uses a frozen RemoteCLIP text encoder to extract text features. These text features are introduced through visual-text interaction to enhance high-level visual features for more reliable coarse matching. CFDM then establishes coarse correspondences based on the enhanced high-level features and refines the matched locations using low-level features. Experimental results on cross-modal remote sensing images demonstrate the effectiveness of TAR, which achieves stronger matching performance than several state-of-the-art methods and yields significant gains under large geometric deformations.

图像配准遥感跨模态文本增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。