解决个性化图像生成中模型失真问题,保持原能力同时精准贴合用户提示。
Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional Drift

- 用基于利普希茨的正则化约束参数更新,防止分布漂移。
- 在多个模型上实现更高视觉保真度与提示遵循率。
- 计算高效,适合资源受限场景,适合需定制化生成的开发者。
个性化文本到图像扩散模型需从少量参考图中引入新视觉概念,同时保留原始生成能力。然而,现有方法常导致过拟合,模型忽略用户提示而仅复制参考图。我们发现根本原因在于个人化目标(主体保真度与文本对齐)与训练目标不匹配:现有方法未显式保留预训练模型的输出分布,引发分布漂移,损害多样性与连贯性。为此,我们提出一种基于利普希茨的正则化目标,在个性化过程中约束参数更新,确保与原始分布偏差有界。该方法在保持预训练模型行为一致性的同时,精准适配新概念。此外,该方法比常用但高耗能的采样技术更高效。在多种扩散模型架构上的大量实验表明,本方法在定量指标与定性评估中均表现优异,持续提升视觉保真度与提示遵循性。我们通过消融研究与可视化分析进一步验证了其有效性。
原文摘要 · Abstract (English)
Personalizing text-to-image diffusion models involves integrating novel visual concepts from a small set of reference images while retaining the model's original generative capabilities. However, this process often leads to overfitting, where the model ignores the user's prompt and merely replicates the reference images. We attribute this issue to a fundamental misalignment between the true goals of personalization, which are subject fidelity and text alignment, and the training objectives of existing methods that fail to enforce both objectives simultaneously. Specifically, prior approaches often overlook the need to explicitly preserve the pretrained model's output distribution, resulting in distributional drift that undermines diversity and coherence. To resolve these challenges, we introduce a Lipschitz-based regularization objective that constrains parameter updates during personalization, ensuring bounded deviation from the original distribution. This promotes consistency with the pretrained model's behavior while enabling accurate adaptation to new concepts. Furthermore, our method offers a computationally efficient alternative to commonly used, resource-intensive sampling techniques. Through extensive experiments across diverse diffusion model architectures, we demonstrate that our approach achieves superior performance in both quantitative metrics and qualitative evaluations, consistently excelling in visual fidelity and prompt adherence. We further support these findings with comprehensive analyses, including ablation studies and visualizations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。