arXiv:2506.06483cs.GRcs.AI2025-06CVPR被引 2

通过一致性正则化提升扩散模型对特定主体的生成稳定性与多样性

Noise Consistency Regularization for Improved Subject-Driven Image Synthesis

  • 引入先验与主体双重一致性正则化,约束噪声预测保持一致
  • 在保留主体特征的同时,显著提升背景多样性和图像质量
  • 适合需要高保真主体生成的个性化图像合成场景

微调Stable Diffusion可实现以特定主体为中心的图像生成,但现有方法存在两个关键问题:欠拟合导致主体身份识别不可靠,过拟合则使模型记忆主体图像并降低背景多样性。为此,我们提出两种辅助一致性损失用于扩散模型微调:一是先验一致性正则化损失,确保非主体图像的预测扩散噪声与预训练模型一致,提升生成保真度;二是主体一致性正则化损失,增强模型对乘性噪声调制潜在编码的鲁棒性,有助于在保持主体身份的同时提升多样性。实验表明,加入这些损失后,不仅有效维持主体身份,还显著提升图像多样性,优于DreamBooth在CLIP分数、背景变化和整体视觉质量上的表现。

原文摘要 · Abstract (English)

Fine-tuning Stable Diffusion enables subject-driven image synthesis by adapting the model to generate images containing specific subjects. However, existing fine-tuning methods suffer from two key issues: underfitting, where the model fails to reliably capture subject identity, and overfitting, where it memorizes the subject image and reduces background diversity. To address these challenges, we propose two auxiliary consistency losses for diffusion fine-tuning. First, a prior consistency regularization loss ensures that the predicted diffusion noise for prior (non-subject) images remains consistent with that of the pretrained model, improving fidelity. Second, a subject consistency regularization loss enhances the fine-tuned model's robustness to multiplicative noise modulated latent code, helping to preserve subject identity while improving diversity. Our experimental results demonstrate that incorporating these losses into fine-tuning not only preserves subject identity but also enhances image diversity, outperforming DreamBooth in terms of CLIP scores, background variation, and overall visual quality.

图像生成扩散模型微调一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。