提出新方法让大模型对齐更灵活,无需成对数据也能高效训练。
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
- 引入不确定性锚点函数,缓解标注数据不完整的问题。
- 在无配对数据下仍能取得与传统方法相当的性能。
- 适合数据稀缺或标注成本高的对齐场景,提升训练鲁棒性。
离线偏好优化方法在大语言模型对齐中效率较高。以直接偏好优化(DPO)为代表的主流方法因奖励建模高效而受到青睐,但通常依赖布拉德利-特里(BT)模型,该模型需满足成对训练数据、模型分布不变、人类理性等关键假设。为突破这些限制,本文提出通用框架自适应偏好优化与效用锚点(UAPO),通过引入锚定函数估计偏好数据标注带来的不确定性。该方法可支持无配对数据训练,显著提升数据利用效率;同时锚点设计增强了训练过程的鲁棒性。实验表明,UAPO在无需严格依赖数据配对的情况下仍能取得竞争力结果,为更灵活高效的偏好优化开辟新路径。
原文摘要 · Abstract (English)
Offline preference optimization methods are efficient for large language models (LLMs) alignment. Direct Preference optimization (DPO)-like learning, one of the most popular approaches, stands out for its efficiency in reward modeling. However, these methods typically follow the convention to use Bradley-Terry (BT) reward modeling that faces several critical assumptions, including the requirement for pairwise training data, model distribution shifting, human rationality assumption, etc. To address these limitations, we propose a general framework for offline preference optimization methods, Adaptive Preference Optimization with Utility Anchor (UAPO), which introduces an anchoring function to estimate the uncertainties brought from preference data annotation. Our method enables training even in scenarios where the data is unpaired, significantly enhancing data utilization efficiency. Moreover, the anchor design makes UAPO more robust in the training process. Experimental results demonstrate that UAPO achieves competitive outcomes without the strict dependency on data pairing, paving the way for more flexible and effective preference optimization methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。