用神经网络提升非线性场景下人类偏好反馈的收集效率
Active Human Feedback Collection via Neural Contextual Dueling Bandits
- 基于神经上下文对决强化学习框架,适配非线性奖励函数
- 理论证明反馈数据越多,策略偏差下降速度为次线性
- 适合需要高效获取人类偏好的推荐系统与大模型对齐任务
收集人类偏好反馈成本高昂,现有方法多假设奖励函数为线性,这在在线推荐和大模型对齐等真实场景中不成立。为此,本文提出神经上下文对决强化学习算法 Neural-ADB,可在非线性潜在奖励函数下高效收集人类偏好反馈。理论上,当偏好反馈服从 Bradley-Terry-Luce 模型时,Neural-ADB 学习到的策略最坏次优差距随偏好数据集规模增加呈次线性下降。实验结果进一步验证了该方法的有效性。
原文摘要 · Abstract (English)
Collecting human preference feedback is often expensive, leading recent works to develop principled algorithms to select them more efficiently. However, these works assume that the underlying reward function is linear, an assumption that does not hold in many real-life applications, such as online recommendation and LLM alignment. To address this limitation, we propose Neural-ADB, an algorithm based on the neural contextual dueling bandit framework that provides a principled and practical method for collecting human preference feedback when the underlying latent reward function is non-linear. We theoretically show that when preference feedback follows the Bradley-Terry-Luce model, the worst sub-optimality gap of the policy learned by Neural-ADB decreases at a sub-linear rate as the preference dataset increases. Our experimental results on preference datasets further corroborate the effectiveness of Neural-ADB.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。