用上下文解释提升语言模型判断力,让奖励模型更稳定可靠。
Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
- 将奖励建模视为自然语言推理任务,利用掩码语言模型捕捉上下文解释。
- 在NLI任务中,结合解释的掩码模型性能显著优于传统自回归模型。
- 新框架ESFP-RM适用于人类和分布外反馈,提升奖励信号稳定性。
基于语言模型的评判奖励建模(如生成式奖励模型)已成功实现从人工智能反馈中强化学习(RLAIF)的高效与可扩展性。为进一步推进该范式,本文提出核心洞察:此类奖励建模与自然语言推理(NLI)具有本质形式一致性。这一视角指向构建更优奖励模型的关键路径——扩展模型的理解边界。在NLI任务上的探索性实验表明,融合上下文解释的槽位预测掩码语言模型(MLMs)相比主流自回归模型表现显著更优。基于此发现,本文提出两阶段语言模型评判奖励模型ESFP-RM,采用基于解释的槽位框架进行预测,充分释放MLMs优势。大量实验显示,在从人类反馈强化学习(RLHF)及分布外(OOD)场景中,ESFP-RM相较生成式奖励模型提供更稳定、泛化性更强的奖励信号。
原文摘要 · Abstract (English)
The emergence of LM-based judging reward modeling, represented by generative reward models, has successfully made reinforcement learning from AI feedback (RLAIF) efficient and scalable. To further advance this paradigm, we propose a core insight: this form of reward modeling shares fundamental formal consistency with natural language inference (NLI), a core task in natural language understanding. This reframed perspective points to a key path for building superior reward models: scaling the model's comprehension boundaries. Pursuing this path, exploratory experiments on NLI tasks demonstrate that the slot prediction masked language models (MLMs) incorporating contextual explanations achieve significantly better performance compared to mainstream autoregressive models. Based on this key finding, we propose ESFP-RM, a two-stage LM-based judging reward model that utilizes an explanation based slot framework for prediction to fully leverage the advantages of MLMs. Extensive experiments demonstrate that in both reinforcement learning from human feedback (RLHF) and out-of-distribution (OOD) scenarios, the ESFP-RM framework delivers more stable and generalizable reward signals compared to generative reward models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。