揭示深度神经网络奖励建模的理论保证,强调人类判断清晰性的重要性。
Learning Guarantee of Reward Modeling Using Deep Neural Networks
- 基于成对比较数据,建立非参数下深度奖励估计器的非渐近误差界。
- 引入边缘条件使误差界更紧,解释了强化学习中人类反馈的有效性。
- 适用于多种模型与算法,核心依赖高质量成对数据而非具体方法。
本文研究使用深度神经网络进行奖励建模的理论学习问题,基于成对比较数据。在非参数设定下,建立了深度奖励估计器的新型非渐近遗憾界,该界显式依赖于网络结构。为强调人类信念清晰性的重要性,引入一种边缘型条件,假设最优动作在成对比较中的胜率显著偏离1/2。该条件使遗憾界更紧,验证了从人类反馈中强化学习的实证效率,并凸显清晰人类信念在其成功中的作用。值得注意的是,此改进源于边缘条件所隐含的高质量成对比较数据,与具体估计器无关,可推广至多种学习算法和模型。
原文摘要 · Abstract (English)
In this work, we study the learning theory of reward modeling with pairwise comparison data using deep neural networks. We establish a novel non-asymptotic regret bound for deep reward estimators in a non-parametric setting, which depends explicitly on the network architecture. Furthermore, to underscore the critical importance of clear human beliefs, we introduce a margin-type condition that assumes the conditional winning probability of the optimal action in pairwise comparisons is significantly distanced from 1/2. This condition enables a sharper regret bound, which substantiates the empirical efficiency of Reinforcement Learning from Human Feedback and highlights clear human beliefs in its success. Notably, this improvement stems from high-quality pairwise comparison data implied by the margin-type condition, is independent of the specific estimators used, and thus applies to various learning algorithms and models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。