用可解释概念构建奖励模型,提升透明度和标注效率。
Interpretable Reward Modeling with Active Concept Bottlenecks
- 通过选择性标注概念实现可解释的偏好学习
- 在低监督下仅用少量标注即达更高准确率
- 适合需要透明可控奖励模型的研究与应用
我们提出概念瓶颈奖励模型(CB-RM),一种通过选择性概念标注实现可解释偏好学习的奖励建模框架。与依赖不透明奖励函数的标准强化学习人类反馈(RLHF)方法不同,CB-RM将奖励预测分解为人类可理解的概念。为提升该框架在低监督场景下的效率,我们形式化了一种主动学习策略,动态获取最具信息量的概念标签。基于期望信息增益设计的获取函数显著加速了概念学习过程,且不牺牲偏好准确性。在UltraFeedback数据集上的评估表明,本方法在可解释性和样本效率方面均优于基线,标志着向更透明、可审计、与人类对齐的奖励模型迈进一步。
原文摘要 · Abstract (English)
We introduce Concept Bottleneck Reward Models (CB-RM), a reward modeling framework that enables interpretable preference learning through selective concept annotation. Unlike standard RLHF methods that rely on opaque reward functions, CB-RM decomposes reward prediction into human-interpretable concepts. To make this framework efficient in low-supervision settings, we formalize an active learning strategy that dynamically acquires the most informative concept labels. We propose an acquisition function based on Expected Information Gain and show that it significantly accelerates concept learning without compromising preference accuracy. Evaluated on the UltraFeedback dataset, our method outperforms baselines in interpretability and sample efficiency, marking a step towards more transparent, auditable, and human-aligned reward models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。