arXiv:2412.08464cs.CV2024-12被引 13

提升遥感图像生成中前景与背景的一致性,让合成更真实。

CC-Diff: Enhancing Contextual Coherence in Remote Sensing Image Synthesis

  • 用双重重采样器和上下文桥机制捕捉前景与背景的依赖关系。
  • 在遥感图像生成上优于主流方法,检测任务准确率提升超2个百分点。
  • 适合需要高质量合成图像的遥感分析、目标检测研究者使用。

现有自然场景图像生成方法多关注前景控制,常将背景简化为单调纹理,忽略前景与背景间的内在关联,导致遥感场景下合成结果不一致且失真。本文提出基于扩散模型的CC-Diff方法,增强遥感图像生成中的上下文一致性。设计新型双重重采样器,内置“上下文桥”以显式捕捉前景与背景间的复杂依赖;同时在生成背景特征时引入前景感知注意力机制,强化二者联系,提升合成场景合理性。大量实验表明,CC-Diff在关键质量指标上超越当前最优方法,在遥感领域表现优异,并可泛化至自然图像生成。尤为显著的是,该方法具有高可训练性,在DOTA数据集上提升1.83 mAP,CO CO数据集上提升2.25 mAP。

原文摘要 · Abstract (English)

Existing image synthesis methods for natural scenes focus primarily on foreground control, often reducing the background to simplistic textures. Consequently, these approaches tend to overlook the intrinsic correlation between foreground and background, which may lead to incoherent and unrealistic synthesis results in remote sensing (RS) scenarios. In this paper, we introduce CC-Diff, a $\underline{\textbf{Diff}}$usion Model-based approach for RS image generation with enhanced $\underline{\textbf{C}}$ontext $\underline{\textbf{C}}$oherence. Specifically, we propose a novel Dual Re-sampler for feature extraction, with a built-in `Context Bridge' to explicitly capture the intricate interdependency between foreground and background. Moreover, we reinforce their connection by employing a foreground-aware attention mechanism during the generation of background features, thereby enhancing the plausibility of the synthesized context. Extensive experiments show that CC-Diff outperforms state-of-the-art methods across critical quality metrics, excelling in the RS domain and effectively generalizing to natural images. Remarkably, CC-Diff also shows high trainability, boosting detection accuracy by 1.83 mAP on DOTA and 2.25 mAP on the COCO benchmark.

遥感图像扩散模型上下文一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。