解决低光水下图像生成中背景干扰问题,提升布局对齐与生成质量。
FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation
- 通过频率感知解耦机制,分离前景与背景的频域特征。
- 在5个退化场景上生成质量超越现有方法,布局对齐度显著提升。
- 适合需要高保真退化图像生成的研究者或工业应用开发者。
布局到图像(L2I)生成在自然场景中表现良好,但在低光、水下等退化场景中面临生成保真度低和布局对齐弱的问题。我们将其归因于退化条件下存在的“上下文幻觉困境”,即前景被主导频域分布的背景所掩盖。为此,本文提出频率启发的上下文解耦生成框架(FICGen),将退化图像的频率知识迁移到潜在扩散空间,通过频域感知引导实现退化实例及其环境的高质量渲染。具体而言,FICGen包含两步:首先引入可学习的双查询机制,搭配专用频域重采样器,从训练集退化样本中提取上下文频域原型;其次采用视觉-频率增强注意力机制,将频域原型注入生成过程。为缓解上下文幻觉与特征泄露,设计实例一致性图以调控个体实例与周围环境的潜在空间解耦,并结合自适应空间-频域聚合模块重建混合退化表示。在涵盖严重低光至轻度模糊等多样化退化场景的5个基准测试上,实验表明FICGen在生成保真度、布局对齐及下游辅助可训练性方面均持续优于现有L2I方法。
原文摘要 · Abstract (English)
Layout-to-image (L2I) generation has exhibited promising results in natural domains, but suffers from limited generative fidelity and weak alignment with user-provided layouts when applied to degraded scenes (i.e., low-light, underwater). We primarily attribute these limitations to the "contextual illusion dilemma" in degraded conditions, where foreground instances are overwhelmed by context-dominant frequency distributions. Motivated by this, our paper proposes a new Frequency-Inspired Contextual Disentanglement Generative (FICGen) paradigm, which seeks to transfer frequency knowledge of degraded images into the latent diffusion space, thereby facilitating the rendering of degraded instances and their surroundings via contextual frequency-aware guidance. To be specific, FICGen consists of two major steps. Firstly, we introduce a learnable dual-query mechanism, each paired with a dedicated frequency resampler, to extract contextual frequency prototypes from pre-collected degraded exemplars in the training set. Secondly, a visual-frequency enhanced attention is employed to inject frequency prototypes into the degraded generation process. To alleviate the contextual illusion and attribute leakage, an instance coherence map is developed to regulate latent-space disentanglement between individual instances and their surroundings, coupled with an adaptive spatial-frequency aggregation module to reconstruct spatial-frequency mixed degraded representations. Extensive experiments on 5 benchmarks involving a variety of degraded scenarios-from severe low-light to mild blur-demonstrate that FICGen consistently surpasses existing L2I methods in terms of generative fidelity, alignment and downstream auxiliary trainability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。