arXiv:2602.10662cs.CV2026-02

通过动态调制频率成分,实现文本驱动图像生成中结构稳定与语义修改的平衡。

Dynamic Frequency Modulation for Controllable Text-driven Image Generation

  • 从频域角度分析生成过程,发现低频成分主导早期结构构建
  • 提出无需训练的动态衰减加权方法,直接操控噪声潜在变量
  • 避免人工选择特征图,显著提升结构一致性与语义修改精度

文本引导的扩散模型已确立以迭代优化文本提示为驱动的新图像生成范式。然而,修改原始提示以实现预期语义调整常导致全局结构意外变化,违背用户意图。现有方法依赖经验选择特征图进行干预,性能高度依赖选取得当,稳定性不足。本文从频率视角出发,分析噪声潜在变量的频谱对生成过程中层级结构框架与细粒度纹理逐步浮现的影响。研究发现,低频成分在生成初期主要负责结构框架建立,其影响随时间动态衰减,高频频段则主导细粒度纹理合成。基于此,提出一种无需训练的频率调制方法,采用具有动态衰减特性的频域加权函数。该方法在保持结构框架一致性的同时,支持精准的语义修改。通过直接操纵噪声潜在变量,避免了内部特征图的经验选择。大量实验表明,该方法显著优于当前最先进方法,在结构保留与语义更新之间实现有效平衡。

原文摘要 · Abstract (English)

The success of text-guided diffusion models has established a new image generation paradigm driven by the iterative refinement of text prompts. However, modifying the original text prompt to achieve the expected semantic adjustments often results in unintended global structure changes that disrupt user intent. Existing methods rely on empirical feature map selection for intervention, whose performance heavily depends on appropriate selection, leading to suboptimal stability. This paper tries to solve the aforementioned problem from a frequency perspective and analyzes the impact of the frequency spectrum of noisy latent variables on the hierarchical emergence of the structure framework and fine-grained textures during the generation process. We find that lower-frequency components are primarily responsible for establishing the structure framework in the early generation stage. Their influence diminishes over time, giving way to higher-frequency components that synthesize fine-grained textures. In light of this, we propose a training-free frequency modulation method utilizing a frequency-dependent weighting function with dynamic decay. This method maintains the structure framework consistency while permitting targeted semantic modifications. By directly manipulating the noisy latent variable, the proposed method avoids the empirical selection of internal feature maps. Extensive experiments demonstrate that the proposed method significantly outperforms current state-of-the-art methods, achieving an effective balance between preserving structure and enabling semantic updates.

图像生成扩散模型频率调制可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。