通过对比学习提升零样本主体生成的细节保真度
Negative-Guided Subject Fidelity Optimization for Zero-Shot Subject-Driven Generation
- 引入合成负样本,通过正负对比优化主体细节
- 在中间扩散步骤重加权,聚焦主体特征生成阶段
- 无需人工标注,适合需要高保真主体生成的应用
我们提出主体保真度优化(SFO),一种用于零样本主体驱动生成的新颖比较学习框架,旨在提升主体保真度。现有监督微调方法仅依赖正样本,且沿用预训练阶段的扩散损失,难以捕捉细粒度主体细节。为解决此问题,SFO引入额外的合成负样本,并通过成对比较显式引导模型偏好正样本而非负样本。针对负样本,我们提出条件退化负采样(CDNS),通过引入可控退化自动生成适用于主体驱动生成的合成负样本,强调主体保真度与文本对齐,且无需昂贵的人工标注。此外,我们对扩散时间步进行重加权,使微调聚焦于主体细节显现的中间阶段。大量实验表明,结合CDNS的SFO在主体驱动生成基准上显著优于近期强基线,在主体保真度和文本对齐方面均表现更优。
原文摘要 · Abstract (English)
We present Subject Fidelity Optimization (SFO), a novel comparative learning framework for zero-shot subject-driven generation that enhances subject fidelity. Existing supervised fine-tuning methods, which rely only on positive targets and use the diffusion loss as in the pre-training stage, often fail to capture fine-grained subject details. To address this, SFO introduces additional synthetic negative targets and explicitly guides the model to favor positives over negatives through pairwise comparison. For negative targets, we propose Condition-Degradation Negative Sampling (CDNS), which automatically produces synthetic negatives tailored for subject-driven generation by introducing controlled degradations that emphasize subject fidelity and text alignment without expensive human annotations. Moreover, we reweight the diffusion timesteps to focus fine-tuning on intermediate steps where subject details emerge. Extensive experiments demonstrate that SFO with CDNS significantly outperforms recent strong baselines in terms of both subject fidelity and text alignment on a subject-driven generation benchmark. Project page: https://subjectfidelityoptimization.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。