arXiv:2602.06355cs.CVcs.AI2026-02

通过分离目标区域提升图像生成偏好优化效率

Di3PO - Diptych Diffusion DPO for Targeted Improvements in Image Generation

  • 构造正负样本时仅改变目标区域,保持上下文稳定
  • 在文本渲染任务上优于SFT和DPO基线方法
  • 避免无关像素变异,提升训练效率与可控性

现有文本到图像扩散模型的偏好微调方法通常依赖计算成本高昂的图像生成步骤来构建正负样本对。这些方法常产生差异不明显、采样过滤代价高或在无关像素区域存在显著方差的训练对,从而降低训练效率。为此,我们提出「Di3PO」,一种新型正负样本构建方法,能在偏好微调中精准隔离需改进的特定区域,同时保持图像周围上下文稳定。我们在扩散模型的文本渲染这一挑战性任务上验证了该方法的有效性,结果显示其性能优于SFT和DPO基线方法。

原文摘要 · Abstract (English)

Existing methods for preference tuning of text-to-image (T2I) diffusion models often rely on computationally expensive generation steps to create positive and negative pairs of images. These approaches frequently yield training pairs that either lack meaningful differences, are expensive to sample and filter, or exhibit significant variance in irrelevant pixel regions, thereby degrading training efficiency. To address these limitations, we introduce "Di3PO", a novel method for constructing positive and negative pairs that isolates specific regions targeted for improvement during preference tuning, while keeping the surrounding context in the image stable. We demonstrate the efficacy of our approach by applying it to the challenging task of text rendering in diffusion models, showcasing improvements over baseline methods of SFT and DPO.

图像生成扩散模型偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。