用隐式噪声引导提升文本生成图像的控制力与质量。
The Silent Assistant: NoiseQuery as Implicit Guidance for Goal-Driven Image Generation
- 引入隐式噪声查询,补充文本提示实现更精准控制。
- 无需微调即可跨模型通用,显著提升生成质量与细节控制。
- 兼容现有流程,计算开销极低,适合实际应用。
本文提出NoiseQuery,一种用于增强多目标驱动文本到图像生成的新型噪声初始化方法。通过利用对齐的高斯噪声作为隐式引导,补充文本提示等显式输入,以提升生成质量和可控性。不同于针对特定模型设计的噪声优化方法,该方法基于对扩散模型中通用有限步噪声调度机制的根本分析,实现无需调优的跨架构泛化。其模型无关特性使得可构建适用于多种文本到图像模型及增强技术的可复用噪声库,构成高效生成的基础层。大量实验表明,NoiseQuery不仅在高层次语义上表现优异,还能有效控制低层次视觉属性(通常难以仅靠文本指定),且能无缝集成至现有工作流,计算开销极小。
原文摘要 · Abstract (English)
In this work, we introduce NoiseQuery as a novel method for enhanced noise initialization in versatile goal-driven text-to-image (T2I) generation. Specifically, we propose to leverage an aligned Gaussian noise as implicit guidance to complement explicit user-defined inputs, such as text prompts, for better generation quality and controllability. Unlike existing noise optimization methods designed for specific models, our approach is grounded in a fundamental examination of the generic finite-step noise scheduler design in diffusion formulation, allowing better generalization across different diffusion-based architectures in a tuning-free manner. This model-agnostic nature allows us to construct a reusable noise library compatible with multiple T2I models and enhancement techniques, serving as a foundational layer for more effective generation. Extensive experiments demonstrate that NoiseQuery enables fine-grained control and yields significant performance boosts not only over high-level semantics but also over low-level visual attributes, which are typically difficult to specify through text alone, with seamless integration into current workflows with minimal computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。