arXiv:2511.12867cs.AIcs.LG2025-11AAAI

用偏好优化自我迭代训练大模型,无需大量人工标注。

Bootstrapping LLMs via Preference-Based Policy Optimization

  • 构建主策略与奖励模型的博弈框架,动态优化。
  • 在五个基准上超越现有最优方法,性能持续提升。
  • 适合追求低成本对齐大模型的科研与工程人员。

通过基于偏好的策略优化实现大语言模型的自举训练,为在不依赖大量人工标注的情况下对齐模型行为与人类偏好提供了新方向。本文提出一种新型偏好优化(PbPO)框架,将学习过程建模为主策略与奖励模型(RM)之间的最小最大博弈。奖励模型被约束于由偏好数据导出的置信集内,以确保可靠利用。我们设计了迭代在线算法,通过引导对不断演化的策略进行探索,主动收集偏好数据,实现策略与奖励模型的持续自我改进。理论分析表明,该方法在序列级和词元级奖励模型设定下均具有高概率的遗憾界,证明了其在大模型自举中的有效性。在五个基准上的大量实验显示,该方法始终优于现有的先进偏好优化技术。

原文摘要 · Abstract (English)

Bootstrapping large language models (LLMs) through preference-based policy optimization offers a promising direction for aligning model behavior with human preferences without relying on extensive manual annotations. In this work, we propose a novel preference-based policy optimization (PbPO) framework that formulates the learning process as a min-max game between the main policy and a reward model (RM). The RM is constrained within a confidence set derived from preference data to ensure reliable exploitation. Our iterative online algorithm actively collects preference data through guided exploration of the evolving policy, enabling continual self-improvement of both the policy and the RM. We provide theoretical guarantees for our method, establishing high-probability regret bounds for both settings with sequence-level RM and token-level RM, demonstrating its effectiveness in bootstrapping LLMs. Extensive experiments on five benchmarks show that our approach consistently outperforms existing state-of-the-art preference optimization techniques.

大模型对齐偏好优化自举训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。