提出新模型解决遥感图像超分辨中参考图依赖过强或过弱的问题。
Learning to Balance: Decoupled Siamese Diffusion Transformer for Reference-Based Remote Sensing Image Super-Resolution

- 解耦注意力机制,让低分辨率结构与参考图纹理独立作用于噪声隐变量。
- 在多个数据集上实现更优的定量指标和视觉细节恢复效果。
- 适合需要高保真遥感图像重建的研究者与应用开发者。
基于扩散的方法在大尺度遥感图像超分辨率任务中展现出巨大潜力,尤其在参考图像引导的超分辨率(RefSR)中,高分辨率参考图像能提供关键的细粒度纹理先验。然而,现有方法常面临对参考信息过度依赖(导致纹理伪影)或利用不足(细节恢复不充分)的权衡问题。为此,我们提出DS-DiT:一种解耦式孪生扩散变换器,通过在注意力机制中解耦低分辨率(LR)与参考(Ref)条件之间的交互,使LR结构先验与Ref纹理信息可独立与噪声隐变量交互,有效缓解两者间的竞争。为弥补全局注意力的局部建模局限,引入分块加权(PLW)模块,自适应调节条件信息融合。此外,孪生架构支持推理时自动引导策略,利用强/弱参考条件下的预测差异提升生成质量,无需额外训练。多数据集、多放大倍数的实验结果表明,DS-DiT在定量指标和视觉保真度上均优于现有方法。
原文摘要 · Abstract (English)
Diffusion-based methods demonstrate significant potential for remote sensing image super-resolution at large scaling factors, particularly in reference-based super-resolution (RefSR), where high-resolution reference images provide critical fine-grained texture priors. However, existing methods often suffer from a trade-off between over-reliance on reference information, which leads to texture artifacts, and under-utilization of such information, which results in insufficient detail recovery. To address these issues, we propose DS-DiT, a Decoupled Siamese Diffusion Transformer that decouples the interaction between low-resolution (LR) and reference (Ref) conditions within the attention mechanism. By allowing LR structural priors and Ref texture information to independently interact with the noisy latent, the framework effectively mitigates competition between the two conditional sources. To further compensate for the limited local modeling ability of global attention, we introduce a Patch-Level Weighting (PLW) module that adaptively modulates the fusion of conditional sources. In addition, the siamese architecture enables an inference-time autoguidance strategy that exploits the prediction discrepancy between strong and weak Ref conditions to improve generation quality without additional training. Experimental results across multiple datasets and scaling factors show that DS-DiT outperforms existing methods in both quantitative metrics and visual fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。