用CLIP做语言监督,实现无需真值的全分辨率无监督遥感图像融合。
CLIPPan: Adapting CLIP as A Supervisor for Unsupervised Pansharpening
- 用轻量微调让CLIP理解多光谱、全色及融合图像特征
- 通过语义提示构建语言约束损失,提升融合图像保真度
- 适用于真实遥感数据,适合遥感图像处理研究者
尽管监督式全色锐化神经网络取得了显著进展,但其在分辨率域适应方面仍面临挑战,主要源于模拟低分辨率训练数据与真实世界全分辨率场景之间的固有差异。为弥合这一差距,我们提出一种无监督全分辨率全色锐化框架CLIPPan,将视觉-语言模型CLIP作为监督信号。由于CLIP本身偏向自然图像且对全色锐化任务理解有限,我们首先设计了一种轻量级微调流程,使CLIP能够识别低分辨率多光谱、全色及高分辨率多光谱图像,并理解融合过程。在此基础上,我们提出一种新型损失函数,整合语义语言约束,将图像级融合结果与标准文本描述(如Wald或Khan描述)对齐,从而在无真实标签情况下,利用语言作为强大监督信号引导融合学习。大量实验表明,CLIPPan在多个全色锐化骨干网络上均显著提升光谱与空间保真度,在真实数据集上达到无监督全分辨率全色锐化新基准。
原文摘要 · Abstract (English)
Despite remarkable advancements in supervised pansharpening neural networks, these methods face domain adaptation challenges of resolution due to the intrinsic disparity between simulated reduced-resolution training data and real-world full-resolution scenarios.To bridge this gap, we propose an unsupervised pansharpening framework, CLIPPan, that enables model training at full resolution directly by taking CLIP, a visual-language model, as a supervisor. However, directly applying CLIP to supervise pansharpening remains challenging due to its inherent bias toward natural images and limited understanding of pansharpening tasks. Therefore, we first introduce a lightweight fine-tuning pipeline that adapts CLIP to recognize low-resolution multispectral, panchromatic, and high-resolution multispectral images, as well as to understand the pansharpening process. Then, building on the adapted CLIP, we formulate a novel \textit{loss integrating semantic language constraints}, which aligns image-level fusion transitions with protocol-aligned textual prompts (e.g., Wald's or Khan's descriptions), thus enabling CLIPPan to use language as a powerful supervisory signal and guide fusion learning without ground truth. Extensive experiments demonstrate that CLIPPan consistently improves spectral and spatial fidelity across various pansharpening backbones on real-world datasets, setting a new state of the art for unsupervised full-resolution pansharpening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。