用更通用的散度优化,提升文生图对齐效果与多样性
Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences through $f$-divergence Minimization
- 将文本到图像对齐从仅用KL散度扩展为f-散度框架
- 基于Jensen-Shannon散度的对齐在性能与多样性间平衡最优
- 为实际应用选择合适散度提供了关键依据
直接偏好优化(DPO)已成功将大语言模型对齐方法拓展至文本到图像模型,以匹配人类偏好,引发广泛关注。然而,现有方法仅依赖反向Kullback-Leibler散度进行对齐,忽略了其他散度约束的潜力。本文将文生图对齐范式中的反向KL散度推广至f-散度,旨在提升对齐性能与生成多样性。我们推导了f-散度条件下的通用对齐公式,并从梯度场角度分析不同散度约束对对齐过程的影响。在图像-文本对齐、人类价值对齐及生成多样性三方面进行综合评估,结果表明:基于Jensen-Shannon散度的对齐在三者间取得最佳平衡。散度选择显著影响对齐性能(尤其是人类价值观对齐)与生成多样性的权衡,凸显了在实际应用中选取合适散度的重要性。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) has recently expanded its successful application from aligning large language models (LLMs) to aligning text-to-image models with human preferences, which has generated considerable interest within the community. However, we have observed that these approaches rely solely on minimizing the reverse Kullback-Leibler divergence during alignment process between the fine-tuned model and the reference model, neglecting the incorporation of other divergence constraints. In this study, we focus on extending reverse Kullback-Leibler divergence in the alignment paradigm of text-to-image models to $f$-divergence, which aims to garner better alignment performance as well as good generation diversity. We provide the generalized formula of the alignment paradigm under the $f$-divergence condition and thoroughly analyze the impact of different divergence constraints on alignment process from the perspective of gradient fields. We conduct comprehensive evaluation on image-text alignment performance, human value alignment performance and generation diversity performance under different divergence constraints, and the results indicate that alignment based on Jensen-Shannon divergence achieves the best trade-off among them. The option of divergence employed for aligning text-to-image models significantly impacts the trade-off between alignment performance (especially human value alignment) and generation diversity, which highlights the necessity of selecting an appropriate divergence for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。