提出防作弊的强化学习算法,让人类反馈更真实可靠。
Strategyproof Reinforcement Learning from Human Feedback
- 用悲观中位最大似然法应对反馈欺诈
- 多标注者作假可导致政策严重偏离社会最优
- 适合关注公平性和激励相容的AI系统设计者
我们研究多人标注时可能故意歪曲反馈以影响策略学习的强化学习从人类反馈(RLHF)场景。现有方法,包括近期的多元方法,均不具备策略鲁棒性;即使一个策略性标注者也能导致与社会福利的任意程度偏差。我们证明,在最坏情况下,任何策略鲁棒的RLHF算法性能最多比最优策略差k倍,其中k为标注者数量。这揭示了激励对齐与策略对齐之间的根本权衡。为此,我们提出悲观中位最大似然(Pessimistic Median of MLEs)算法,在合理的策略覆盖假设下,近似策略鲁棒,并随标注者和样本数增加收敛至最优策略。结果适用于上下文老虎机和马尔可夫决策过程。
原文摘要 · Abstract (English)
We study Reinforcement Learning from Human Feedback (RLHF) in settings where multiple labelers may strategically misreport feedback to steer the learned policy toward their own preferences. We show that existing RLHF algorithms, including recent pluralistic methods, are not strategyproof, and that even a single strategic labeler can cause arbitrarily large misalignment with social welfare. Moreover, we prove that, in the worst case, any strategyproof RLHF algorithm must perform $k$-times worse than the optimal policy, where $k$ is the number of labelers. This suggests a fundamental trade-off between incentive alignment (ensuring labelers report truthfully) and policy alignment (maximizing social welfare). To address this, we propose the Pessimistic Median of MLEs algorithm, which, under appropriate policy coverage assumptions, is approximately strategyproof and converges to the optimal policy as the number of labelers and samples increases. Our results apply to both contextual bandits and Markov decision processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。