通过频率分解控制扩散模型生成细节与结构。
FreSca: Scaling in Frequency Space Enhances Diffusion Models
- 将噪声差异分解为高低频成分,分别独立缩放。
- 无需重训练,在多模型上提升生成质量与结构控制。
- 适合需要精细调控图像结构的生成与编辑任务。
潜空间扩散模型(LDMs)在多种图像任务中取得显著成功,但对全局结构与细节实现细粒度、解耦控制仍具挑战。本文系统分析了像素空间、VAE潜空间及LDM内部表示中的频率特性,发现每步t的分类器无关引导产生的“噪声差异”项是极具语义信息且有效的操控目标。基于此,提出FreSca——一种新颖且即插即用的框架,将噪声差异分解为低频与高频成分,通过空间或能量阈值分别施加独立缩放因子。该方法无需模型重训练或架构修改,具备模型与任务无关性。实验证明其在多个架构(如SD3、SDXL)和应用(图像生成、编辑、深度估计、视频合成)中均有效提升生成质量与结构强调能力,为LDM开辟了新的表达控制维度。
原文摘要 · Abstract (English)
Latent diffusion models (LDMs) have achieved remarkable success in a variety of image tasks, yet achieving fine-grained, disentangled control over global structures versus fine details remains challenging. This paper explores frequency-based control within latent diffusion models. We first systematically analyze frequency characteristics across pixel space, VAE latent space, and internal LDM representations. This reveals that the "noise difference" term, derived from classifier-free guidance at each step t, is a uniquely effective and semantically rich target for manipulation. Building on this insight, we introduce FreSca, a novel and plug-and-play framework that decomposes noise difference into low- and high-frequency components and applies independent scaling factors to them via spatial or energy-based cutoffs. Essentially, FreSca operates without any model retraining or architectural change, offering model- and task-agnostic control. We demonstrate its versatility and effectiveness in improving generation quality and structural emphasis on multiple architectures (e.g., SD3, SDXL) and across applications including image generation, editing, depth estimation, and video synthesis, thereby unlocking a new dimension of expressive control within LDMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。