提出新引导方法,让个性化图像生成更平衡真实与细节。
Steering Guidance for Personalized Text-to-Image Diffusion Models
- 用空文本条件弱模型动态调节微调后模型的遗忘程度
- 在不增加计算量的前提下提升文本对齐与主体保真度
- 适合需要精细控制图像生成质量的研究者和开发者
个性化文本到图像扩散模型对适配特定目标概念至关重要,可实现多样化图像生成。然而,仅用少量图像微调会带来固有权衡:既要贴合目标分布(如主体保真度),又要保留原始模型的广泛知识(如文本可编辑性)。现有采样引导方法如无分类器引导(CFG)和自动引导(AG)难以有效引导至平衡空间:CFG过度限制于目标分布,而AG牺牲文本对齐。为此,我们提出个性化引导,一种简单有效的方案,利用一个未学习的弱模型,该模型以空文本提示为条件。此外,我们的方法在推理时通过预训练与微调模型权重插值,动态控制弱模型中的遗忘程度。不同于依赖引导系数的方法,本方法显式将输出导向平衡潜在空间,且无需额外计算开销。实验表明,所提引导能有效提升文本对齐与目标分布保真度,并可无缝集成到多种微调策略中。
原文摘要 · Abstract (English)
Personalizing text-to-image diffusion models is crucial for adapting the pre-trained models to specific target concepts, enabling diverse image generation. However, fine-tuning with few images introduces an inherent trade-off between aligning with the target distribution (e.g., subject fidelity) and preserving the broad knowledge of the original model (e.g., text editability). Existing sampling guidance methods, such as classifier-free guidance (CFG) and autoguidance (AG), fail to effectively guide the output toward well-balanced space: CFG restricts the adaptation to the target distribution, while AG compromises text alignment. To address these limitations, we propose personalization guidance, a simple yet effective method leveraging an unlearned weak model conditioned on a null text prompt. Moreover, our method dynamically controls the extent of unlearning in a weak model through weight interpolation between pre-trained and fine-tuned models during inference. Unlike existing guidance methods, which depend solely on guidance scales, our method explicitly steers the outputs toward a balanced latent space without additional computational overhead. Experimental results demonstrate that our proposed guidance can improve text alignment and target distribution fidelity, integrating seamlessly with various fine-tuning strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。