arXiv:2512.15261cs.CV2025-12AAAI被引 4

用新型架构实现多模态图像融合,提升分辨率并支持零样本增强

MMMamba: A Versatile Cross-Modal In Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement

论文配图:MMMamba: A Versatile Cross-Modal In Context Fusion Framework for Pan-Sharpening and Zero-Shot Image Enhancement
图 1 · 摘自论文原文
  • 基于Mamba架构设计跨模态上下文融合框架,线性计算开销下保持强交互能力
  • 在多个数据集上优于现有最先进方法,尤其在零样本图像增强任务中表现突出
  • 适合需要高效高精度图像融合与跨模态处理的研究者和应用开发者

全色锐化旨在通过融合高分辨率全色(PAN)图像与其对应的低分辨率多光谱(MS)图像,生成高分辨率多光谱(HRMS)图像。为实现有效融合,需充分挖掘两模态间的互补信息。传统基于CNN的方法通常依赖通道拼接与固定卷积算子,难以适应多样的空间与光谱变化。尽管交叉注意力机制可实现全局交互,但其计算效率低,易稀释细粒度对应关系,难以捕捉复杂语义关联。最近的多模态扩散变换器(MMDiT)架构在图像生成与编辑任务中表现出色。与交叉注意力不同,MMDiT采用上下文条件化机制,促进更直接高效的跨模态信息交换。本文提出MMMamba,一种用于全色锐化的跨模态上下文融合框架,具备零样本图像超分辨率的能力。基于Mamba架构,该设计确保线性计算复杂度的同时维持强跨模态交互能力。此外,我们引入新颖的多模态交错扫描(MI)机制,有效促进PAN与MS模态间的信息交换。大量实验表明,该方法在多个任务与基准测试中均显著优于现有最先进技术。

原文摘要 · Abstract (English)

Pan-sharpening aims to generate high-resolution multispectral (HRMS) images by integrating a high-resolution panchromatic (PAN) image with its corresponding low-resolution multispectral (MS) image. To achieve effective fusion, it is crucial to fully exploit the complementary information between the two modalities. Traditional CNN-based methods typically rely on channel-wise concatenation with fixed convolutional operators, which limits their adaptability to diverse spatial and spectral variations. While cross-attention mechanisms enable global interactions, they are computationally inefficient and may dilute fine-grained correspondences, making it difficult to capture complex semantic relationships. Recent advances in the Multimodal Diffusion Transformer (MMDiT) architecture have demonstrated impressive success in image generation and editing tasks. Unlike cross-attention, MMDiT employs in-context conditioning to facilitate more direct and efficient cross-modal information exchange. In this paper, we propose MMMamba, a cross-modal in-context fusion framework for pan-sharpening, with the flexibility to support image super-resolution in a zero-shot manner. Built upon the Mamba architecture, our design ensures linear computational complexity while maintaining strong cross-modal interaction capacity. Furthermore, we introduce a novel multimodal interleaved (MI) scanning mechanism that facilitates effective information exchange between the PAN and MS modalities. Extensive experiments demonstrate the superior performance of our method compared to existing state-of-the-art (SOTA) techniques across multiple tasks and benchmarks.

图像融合跨模态Mamba超分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。