arXiv:2502.02588cs.CV2025-02CVPR被引 44

用多模型奖励校准优化扩散模型,无需人工标注数据。

Calibrated Multi-Preference Optimization for Aligning Diffusion Models

  • 通过计算生成样本的胜率,校准多个奖励模型的偏好。
  • 在GenEval和T2I-Compbench上优于DPO等方法。
  • 适合需要高效对齐文本到图像模型的研究者。

将文本到图像(T2I)扩散模型与偏好优化对齐对人类标注数据集有价值,但人工数据收集成本高昂,限制了可扩展性。使用奖励模型是替代方案,但现有偏好优化方法仅利用成对偏好分布,信息利用率低,且无法推广到多偏好场景,难以处理奖励不一致问题。为此,我们提出校准偏好优化(CaPO),一种无需人工标注数据即可对齐T2I扩散模型的新方法。核心在于通过计算预训练模型生成样本的期望胜率,近似整体偏好。此外,提出基于前沿的成对选择方法,从帕累托前沿中选取样本以有效管理多偏好分布。最后,使用回归损失微调扩散模型,使其匹配所选样本对的校准奖励差异。实验表明,CaPO在单奖励与多奖励设置下均持续优于先前方法,如直接偏好优化(DPO),在GenEval和T2I-Compbench等基准上验证有效。

原文摘要 · Abstract (English)

Aligning text-to-image (T2I) diffusion models with preference optimization is valuable for human-annotated datasets, but the heavy cost of manual data collection limits scalability. Using reward models offers an alternative, however, current preference optimization methods fall short in exploiting the rich information, as they only consider pairwise preference distribution. Furthermore, they lack generalization to multi-preference scenarios and struggle to handle inconsistencies between rewards. To address this, we present Calibrated Preference Optimization (CaPO), a novel method to align T2I diffusion models by incorporating the general preference from multiple reward models without human annotated data. The core of our approach involves a reward calibration method to approximate the general preference by computing the expected win-rate against the samples generated by the pretrained models. Additionally, we propose a frontier-based pair selection method that effectively manages the multi-preference distribution by selecting pairs from Pareto frontiers. Finally, we use regression loss to fine-tune diffusion models to match the difference between calibrated rewards of a selected pair. Experimental results show that CaPO consistently outperforms prior methods, such as Direct Preference Optimization (DPO), in both single and multi-reward settings validated by evaluation on T2I benchmarks, including GenEval and T2I-Compbench.

扩散模型偏好优化多奖励图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。