用信息量选数据,让人工标注更高效
Reviving The Classics: Active Reward Modeling in Large Language Model Alignment
- 基于费舍尔信息量筛选最有价值的对比样本
- 在多个模型和数据集上显著提升标注效率
- 适合需要减少人工标注成本的研究者
从人类偏好构建神经奖励模型是强化学习中人类反馈(RLHF)和大语言模型对齐研究的关键环节。由于人工标注稀缺且成本高昂,如何选择最具信息量的对比样本进行标注,是一个重要而困难的问题。本文提出,理想的对比数据集应平衡表示空间的探索性与中等奖励差异样本的可区分性。技术上,我们引入基于费舍尔信息量的选择策略,借鉴经典实验设计理论,应用于深度神经网络奖励模型的最后线性层。实验证明,该方法在多个开源大语言模型和数据集上表现优异,兼具高计算效率与稳定性,优于深度学习与经典统计领域的其他选择方法。消融实验进一步表明,在主动奖励建模中引入跨提示对比,显著提升标注效率,为RLHF中的标注策略优化提供了新思路。
原文摘要 · Abstract (English)
Building neural reward models from human preferences is a pivotal component in reinforcement learning from human feedback (RLHF) and large language model alignment research. Given the scarcity and high cost of human annotation, how to select the most informative pairs to annotate is an essential yet challenging open problem. In this work, we highlight the insight that an ideal comparison dataset for reward modeling should balance exploration of the representation space and make informative comparisons between pairs with moderate reward differences. Technically, challenges arise in quantifying the two objectives and efficiently prioritizing the comparisons to be annotated. To address this, we propose the Fisher information-based selection strategies, adapt theories from the classical experimental design literature, and apply them to the final linear layer of the deep neural network-based reward modeling tasks. Empirically, our method demonstrates remarkable performance, high computational efficiency, and stability compared to other selection methods from deep learning and classical statistical literature across multiple open-source LLMs and datasets. Further ablation studies reveal that incorporating cross-prompt comparisons in active reward modeling significantly enhances labeling efficiency, shedding light on the potential for improved annotation strategies in RLHF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。