arXiv:2605.00548cs.CVcs.GR2026-05International Conf…被引 1

不训练即可控制生成图像的色彩与结构

Colorful-Noise: Training-Free Low-Frequency Noise Manipulation for Color-Based Conditional Image Generation

论文配图:Colorful-Noise: Training-Free Low-Frequency Noise Manipulation for Color-Based Conditional Image Generation
图 1 · 摘自论文原文
  • 用低频图像先验直接操纵扩散模型的低频噪声
  • 仅改变低频成分就可稳定控制整体色彩和构图
  • 无需训练、开销极小,适合快速风格化生成

文生图扩散模型通过逐步将白高斯噪声转化为自然图像来生成图片。白高斯噪声因无结构特性,能从单一文本提示生成多样输出,但这也导致对特定视觉属性的控制力弱、结果不可预测。本文研究扩散模型输入噪声的特性,发现尽管白高斯噪声各频率能量相当,但低频成分主要决定图像全局结构与色彩分布,高频成分则控制细节。基于此,我们证明仅用低频图像先验对低频噪声进行简单操作,即可有效引导生成过程重建这些低频视觉线索。由此提出一种无需训练、开销极小的方法,能够调控图像整体结构与颜色,同时保留高频成分自由生成细节的能力,实现输出多样性。

原文摘要 · Abstract (English)

Text-to-image diffusion models generate images by gradually converting white Gaussian noise into a natural image. White Gaussian noise is well suited for producing diverse outputs from a single text prompt due to its absence of structure. However, this very property limits control over, and predictability of, specific visual attributes, as the noise is not human-interpretable. In this work, we investigate the characteristics of the input noise in diffusion models. We show that, although all frequencies in white Gaussian noise have comparable statistical energy, low-frequency components primarily determine the images global structure and color composition, while high-frequency components control finer details. Building on this observation, we demonstrate that simple manipulations of the low-frequency noise using low-frequency image priors can effectively condition the generation process to reconstruct these low-frequency visual cues. This allows us to define a simple, training-free method with minimal overhead that steers overall image structure and color, while letting high-frequency components freely emerge as fine details, enabling variability across generated outputs.

图像生成扩散模型噪声控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。