arXiv:2604.05965cs.AI2026-04ACL被引 3

用动态协商打破模型对齐的妥协困局,实现更优多偏好平衡。

Beyond Compromise: Pareto-Lenient Consensus for Efficient Multi-Preference LLM Alignment

  • 引入博弈论框架,允许局部退步以换取整体优势,突破传统优化瓶颈。
  • 在多个数据集上超越基线,提升全局帕累托前沿质量,避免过早收敛。
  • 适合关注多目标对齐、追求模型多样性与鲁棒性的研究者使用。

突破单一偏好范式,将大语言模型与多元人类价值观对齐对稳健部署至关重要。现有多目标偏好对齐(MPA)方法主要依赖静态线性加权或刚性梯度投影来处理冲突,但强制避免矛盾或同步下降常导致提前收敛至局部驻点。这些点虽数学稳定,却因规避短期权衡而牺牲了全局帕累托改进潜力。为破解此僵局,本文提出帕累托宽松共识(PLC),一种将对齐视为动态谈判过程的博弈论框架。与刚性方法不同,PLC通过共识驱动的宽松梯度修正,动态容忍局部退步,只要存在足够大的主导联盟盈余,从而推动优化轨迹跳出局部次优均衡,探索远端帕累托最优前沿。理论分析表明,PLC可实现僵局突破并渐近收敛至帕累托共识均衡。大量实验显示,PLC在固定偏好对齐与全局帕累托前沿质量上均优于基线。本工作凸显了协商驱动对齐作为MPA新方向的潜力。代码已开源:https://anonymous.4open.science/r/aaa-6BB8。

原文摘要 · Abstract (English)

Transcending the single-preference paradigm, aligning LLMs with diverse human values is pivotal for robust deployment. Contemporary Multi-Objective Preference Alignment (MPA) approaches predominantly rely on static linear scalarization or rigid gradient projection to navigate these trade-offs. However, by enforcing strict conflict avoidance or simultaneous descent, these paradigms often prematurely converge to local stationary points. While mathematically stable, these points represent a conservative compromise where the model sacrifices potential global Pareto improvements to avoid transient local trade-offs. To break this deadlock, we propose Pareto-Lenient Consensus (PLC), a game-theoretic framework that reimagines alignment as a dynamic negotiation process. Unlike rigid approaches, PLC introduces consensus-driven lenient gradient rectification, which dynamically tolerates local degradation provided there is a sufficient dominant coalition surplus, thereby empowering the optimization trajectory to escape local suboptimal equilibrium and explore the distal Pareto-optimal frontier. Theoretical analysis validates PLC can facilitate stalemate escape and asymptotically converge to a Pareto consensus equilibrium. Moreover, extensive experiments show that PLC surpasses baselines in both fixed-preference alignment and global Pareto frontier quality. This work highlights the potential of negotiation-driven alignment as a promising avenue for MPA. Our codes are available at https://anonymous.4open.science/r/aaa-6BB8.

多偏好对齐博弈论帕累托优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。