用通道压缩扩散模型实现跨传感器快速高精度图像融合
CC-Pan: Channel-wise Compression based Diffusion for Efficient Pan-Sharpening
- 通过单通道变分自编码器将多光谱图像压缩为紧凑潜在表示
- 在三个数据集上超越现有方法,推理速度提升2-3倍
- 无需针对传感器微调,支持不同传感器间通用融合
近期扩散模型为全色锐化带来了新思路,显著提升了融合精度。然而,多数现有模型在像素空间进行扩散,且需为不同多光谱(MS)传感器训练独立模型,存在推理延迟高、传感器依赖性强的问题。本文提出CC-Pan,一种面向高效全色锐化的跨传感器潜在扩散框架。具体地,CC-Pan训练一个逐波段的单通道变分自编码器(VAE),将高分辨率多光谱(HRMS)图像编码为紧凑的潜在表示,天然支持不同传感器下波段数不同的多光谱图像,并为加速推理奠定基础。随后,通过精心设计的单向与双向交互控制结构,将光谱物理特性、全色图和多光谱图注入扩散主干网络,在潜在空间中实现高精度的空间-光谱融合。此外,在扩散模型中央层引入轻量级基于区域的跨波段注意力(RCBA)模块,强化波段间光谱关联,进一步提升光谱一致性与融合精度。在高分二号、QuickBird和世界视界-3上的大量实验表明,CC-Pan在三个基准上均优于当前最优扩散方法,实现2-3倍推理速度提升,并在未见的世界视界-2传感器上表现出强跨传感器泛化能力,无需任何传感器特定重训练。
原文摘要 · Abstract (English)
Recently, diffusion models have brought novel insights to pan-sharpening and notably boosted fusion precision. However, most existing models perform diffusion in the pixel space and train distinct models for different multispectral (MS) sensors, suffering from high inference latency and sensor-specific limitations. In this paper, we present CC-Pan, a cross-sensor latent diffusion framework for efficient pan-sharpening. Specifically, CC-Pan trains a band-wise single-channel variational autoencoder (VAE) to encode high-resolution multispectral (HRMS) images into compact latent representations, naturally supporting MS images with varying band counts across different sensors and establishing a basis for inference acceleration. Spectral physical properties, along with PAN and MS images, are then injected into the diffusion backbone through carefully designed unidirectional and bidirectional interactive control structures, achieving high-precision spatial--spectral fusion in the latent diffusion process. Furthermore, a lightweight region-based cross-band attention (RCBA) module is incorporated at the central layer of the diffusion model, reinforcing inter-band spectral connections to boost spectral consistency and further elevate fusion precision. Extensive experimental results on GaoFen-2, QuickBird, and WorldView-3 demonstrate that CC-Pan outperforms state-of-the-art diffusion-based methods across all three benchmarks, attains a $2$--$3\times$ inference speedup, and exhibits robust cross-sensor generalization capability on the held-out WorldView-2 sensor without any sensor-specific retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。