arXiv:2605.13155cs.CV2026-05中稿 · ICML

用帕累托前沿引导生成,解决多奖励模型对齐难题。

Pareto-Guided Optimal Transport for Multi-Reward Alignment

论文配图:Pareto-Guided Optimal Transport for Multi-Reward Alignment
图 1 · 摘自论文原文
  • 构建特定提示的帕累托前沿,通过最优传输优化样本分布
  • 在人类评估中胜率近80%,多奖励协同度提升11%
  • 适合追求高质量图像生成与多目标平衡的研究者

文本到图像生成模型在偏好优化上取得显著进展,但实现跨多样奖励模型的鲁棒对齐仍具挑战。现有方法依赖加权求和,调参成本高且难以平衡冲突目标。更严重的是,奖励模型优化易引发奖励作弊——奖励分数上升而生成图像质量下降。我们证明,在异构奖励上限下统一优化全局目标会诱发奖励作弊,弱奖励模型的固有不稳定性加剧此风险。为此,提出帕累托前沿引导最优传输(PG-OT)框架:构建提示相关帕累托前沿,并通过分布感知最优传输将劣质样本映射至前沿。同时设计适用于不同奖励信号特性的在线与离线优化策略。为更严格评估,引入联合主导率(JDR)与联合坍塌率(JCR)作为量化多奖励协同与奖励作弊的指标。实验表明,该方法优于强基线,JDR提升11%,人类评估胜率接近80%。

原文摘要 · Abstract (English)

Text-to-image generation models have achieved remarkable progress in preference optimization, yet achieving robust alignment across diverse reward models remains a significant challenge. Existing multi-reward fusion approaches rely on weighted summation, which is costly to tune and insufficient for balancing conflicting objectives. More critically, optimization with reward models is highly susceptible to reward hacking, where reward scores increase while the perceived quality of generated images deteriorates. We demonstrate that optimizing against a unified global target under heterogeneous reward upper bounds can induce reward hacking, a risk further exacerbated by the inherent instability of weak reward models. To mitigate this, we propose a Pareto Frontier-Guided Optimal Transport (PG-OT) framework. Our method constructs a prompt-specific Pareto frontier and maps dominated samples toward it via distribution-aware optimal transport. Furthermore, we develop both online and offline optimization strategies tailored to diverse reward signal characteristics. To provide a more rigorous assessment, we introduce the Joint Domination Rate (JDR) and Joint Collapse Rate (JCR) as principled metrics to quantify multi-reward synergy and reward hacking. Experimental results show that our approach outperforms strong baselines with an 11% gain in JDR and achieves a near 80% win rate in human evaluations.

文本生成多奖励对齐最优传输图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。