用预训练扩散模型实现更精准的雷达图像转可见光图像
C-DiffSET: Leveraging Latent Diffusion for SAR-to-EO Image Translation with Confidence-Guided Reliable Object Generation
- 基于预训练扩散模型,将雷达与可见光图像映射到同一潜在空间
- 引入置信度引导损失,显著减少因时间差异导致的物体错位伪影
- 在多个数据集上达到当前最优效果,适合遥感图像理解研究者
合成孔径雷达(SAR)图像在云层、昼夜和季节变化下仍能提供稳定的环境覆盖,但其噪声大且结构特征独特,难以解读,尤其对非专业用户。为提升可读性,SAR转可见光(EO)图像翻译(SET)技术应运而生。然而,传统方法在有限的SAR-EO数据集上从头训练,易过拟合。为此,本文提出置信度引导的雷达转可见光翻译框架C-DiffSET,利用在自然图像上大规模预训练的潜在扩散模型(LDM),实现对EO域的有效适应。值得注意的是,我们发现预训练的变分自编码器(VAE)编码器能在不同噪声水平的SAR输入下,仍将SAR与EO图像对齐于同一潜在空间。为进一步提升像素级保真度,提出置信度引导扩散(C-Diff)损失,有效缓解因时间差异造成的物体出现/消失等伪影,增强结构准确性。C-DiffSET在多个数据集上达到当前最优(SOTA)性能,显著优于近期图像到图像翻译及SET方法。
原文摘要 · Abstract (English)
Synthetic Aperture Radar (SAR) imagery provides robust environmental and temporal coverage (e.g., during clouds, seasons, day-night cycles), yet its noise and unique structural patterns pose interpretation challenges, especially for non-experts. SAR-to-EO (Electro-Optical) image translation (SET) has emerged to make SAR images more perceptually interpretable. However, traditional approaches trained from scratch on limited SAR-EO datasets are prone to overfitting. To address these challenges, we introduce Confidence Diffusion for SAR-to-EO Translation, called C-DiffSET, a framework leveraging pretrained Latent Diffusion Model (LDM) extensively trained on natural images, thus enabling effective adaptation to the EO domain. Remarkably, we find that the pretrained VAE encoder aligns SAR and EO images in the same latent space, even with varying noise levels in SAR inputs. To further improve pixel-wise fidelity for SET, we propose a confidence-guided diffusion (C-Diff) loss that mitigates artifacts from temporal discrepancies, such as appearing or disappearing objects, thereby enhancing structural accuracy. C-DiffSET achieves state-of-the-art (SOTA) results on multiple datasets, significantly outperforming the very recent image-to-image translation methods and SET methods with large margins.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。