arXiv:2505.23367cs.CVcs.AI2025-05ICCV被引 6

解决遥感图像融合中的跨模态错位问题,提升细节清晰度与泛化能力。

PAN-Crafter: Learning Modality-Consistent Alignment for PAN-Sharpening

  • 通过自监督机制联合重建多光谱与全色图像,利用全色图高频信息增强对齐。
  • 提出双向跨模态注意力模块,实现纹理与结构的动态自适应对齐。
  • 速度超快、内存占用低,适用于真实卫星数据等复杂场景。

全色-多光谱图像融合旨在将高分辨率全色(PAN)图像与低分辨率多光谱(MS)图像融合,生成高分辨率多光谱(HRMS)图像。然而,由于传感器位置、成像时间及分辨率差异导致的跨模态错位,成为根本性挑战。传统深度学习方法假设像素级精确对齐,依赖逐像素重建损失,在错位情况下易引发光谱失真、双重边缘和模糊。为此,本文提出 PAN-Crafter 框架,显式缓解 PAN 与 MS 模态间的错位问题。核心包含:模态自适应重建(MARs),使单一网络联合重建 HRMS 与 PAN 图像,利用 PAN 的高频细节作为辅助自监督信号;以及跨模态对齐感知注意力(CM3A),双向对齐 MS 纹理与 PAN 结构,实现跨模态特征自适应优化。在多个基准数据集上的实验表明,本方法在所有指标上优于当前最先进方法,推理速度提升 50.11 倍,内存占用减少至 0.63 倍。此外,在未见卫星数据集上表现出强泛化能力,验证其在不同条件下的鲁棒性。

原文摘要 · Abstract (English)

PAN-sharpening aims to fuse high-resolution panchromatic (PAN) images with low-resolution multi-spectral (MS) images to generate high-resolution multi-spectral (HRMS) outputs. However, cross-modality misalignment -- caused by sensor placement, acquisition timing, and resolution disparity -- induces a fundamental challenge. Conventional deep learning methods assume perfect pixel-wise alignment and rely on per-pixel reconstruction losses, leading to spectral distortion, double edges, and blurring when misalignment is present. To address this, we propose PAN-Crafter, a modality-consistent alignment framework that explicitly mitigates the misalignment gap between PAN and MS modalities. At its core, Modality-Adaptive Reconstruction (MARs) enables a single network to jointly reconstruct HRMS and PAN images, leveraging PAN's high-frequency details as auxiliary self-supervision. Additionally, we introduce Cross-Modality Alignment-Aware Attention (CM3A), a novel mechanism that bidirectionally aligns MS texture to PAN structure and vice versa, enabling adaptive feature refinement across modalities. Extensive experiments on multiple benchmark datasets demonstrate that our PAN-Crafter outperforms the most recent state-of-the-art method in all metrics, even with 50.11$\times$ faster inference time and 0.63$\times$ the memory size. Furthermore, it demonstrates strong generalization performance on unseen satellite datasets, showing its robustness across different conditions.

图像融合遥感影像跨模态对齐注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。