arXiv:2504.14177cs.AIcs.CL2025-04被引 1

用在线AI奖励直接优化大模型,比传统方法更高效准确

Direct Advantage Regression: Aligning LLMs with Online AI Reward

  • 通过加权监督微调直接回归AI奖励值
  • 在GPT-4-Turbo和MT-bench上优于OAIF与在线RLHF
  • 无需强化学习,实现更高人类-AI一致性

在线人工智能反馈(OAIF)为对齐语言模型提供了替代强化学习人类反馈(RLHF)的前景,利用在线AI偏好进行对齐。然而,简单地以AI取代人类会剥夺大模型从二元信号之外获取更细粒度监督的机会。本文提出直接优势回归(DAR),一种使用在线AI奖励优化策略改进的简单对齐算法,通过加权监督微调实现。作为无强化学习的方法,DAR保持了与在线强化学习对齐流程的理论一致性,同时显著降低实现复杂度并提升学习效率。实证结果表明,相比AI偏好,AI奖励是一种更优的监督形式,持续获得更高的人类-AI一致性。在GPT-4-Turbo和MT-bench上的评估显示,DAR优于OAIF与在线RLHF基线。

原文摘要 · Abstract (English)

Online AI Feedback (OAIF) presents a promising alternative to Reinforcement Learning from Human Feedback (RLHF) by utilizing online AI preference in aligning language models (LLMs). However, the straightforward replacement of humans with AI deprives LLMs from learning more fine-grained AI supervision beyond binary signals. In this paper, we propose Direct Advantage Regression (DAR), a simple alignment algorithm using online AI reward to optimize policy improvement through weighted supervised fine-tuning. As an RL-free approach, DAR maintains theoretical consistency with online RLHF pipelines while significantly reducing implementation complexity and improving learning efficiency. Our empirical results underscore that AI reward is a better form of AI supervision consistently achieving higher human-AI agreement as opposed to AI preference. Additionally, evaluations using GPT-4-Turbo and MT-bench show that DAR outperforms both OAIF and online RLHF baselines.

大模型对齐在线反馈奖励建模监督微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。