优化人类反馈效率,用新策略提升模型对齐效果
Maximizing the efficiency of human feedback in AI alignment: a comparative analysis
- 提出瑞士轮+信息增益的配对策略,自适应选择反馈样本
- 在有限标注预算下,性能超越经典随机配对方法
- 适合资源受限但追求高质量对齐的研究与工程团队
强化学习中的人类反馈(RLHF)依赖偏好建模来对齐机器学习系统与人类价值观,但当前广泛使用的随机成对采样结合布拉德利-特里模型在标注预算受限时存在统计效率瓶颈。本文探索了偏好推断的替代采样与评估策略,受博弈论、统计学与社会选择理论启发。最佳方法Swiss InfoGain采用瑞士淘汰赛制与代理互信息增益配对规则,在有限标注预算下显著优于所有对比方法,且样本效率更高。即便在高资源场景,也能发现优于布拉德利-特里基线的方案。实验表明,自适应、资源感知的策略可减少冗余,增强鲁棒性,并在偏好学习中带来统计显著提升,凸显在对齐质量与人力投入间取得平衡的重要性。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) relies on preference modeling to align machine learning systems with human values, yet the popular approach of random pair sampling with Bradley-Terry modeling is statistically limited and inefficient under constrained annotation budgets. In this work, we explore alternative sampling and evaluation strategies for preference inference in RLHF, drawing inspiration from areas such as game theory, statistics, and social choice theory. Our best-performing method, Swiss InfoGain, employs a Swiss tournament system with a proxy mutual-information-gain pairing rule, which significantly outperforms all other methods in constrained annotation budgets while also being more sample-efficient. Even in high-resource settings, we can identify superior alternatives to the Bradley-Terry baseline. Our experiments demonstrate that adaptive, resource-aware strategies reduce redundancy, enhance robustness, and yield statistically significant improvements in preference learning, highlighting the importance of balancing alignment quality with human workload in RLHF pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。