arXiv:2502.08922cs.AI2025-02被引 12

提升大模型内部奖励的一致性,让自评更可靠。

Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models

  • 用多个内部奖励模型投票,通过惩罚不一致预测提升可靠性
  • 仅使用一致预测数据优化,使对齐效果显著优于基线
  • 适合追求高质量自对齐的大模型研究者

将大语言模型(LLM)与人类偏好对齐是其实现真实应用的关键。近期的自奖励语言模型表明,LLM 可利用其内部奖励模型(如 LLM-as-a-Judge)生成偏好数据,无需昂贵的人工标注即可提升对齐性能。然而,我们发现同一模型内的不同内部奖励模型常产生不一致的偏好判断。这种不一致性影响自生成偏好数据的可信度,阻碍整体对齐效果,并凸显了确保对齐可靠性和一致性的必要性。为此,我们提出自一致性内部奖励(SCIR)框架,在训练中收集多个预定义内部奖励模型的偏好预测,通过不一致惩罚机制强制一致性与置信度,从而提升内部奖励模型的可靠性。我们仅选择一致预测的数据用于偏好优化,保障数据质量。采用自一致性内部奖励后,方法显著提升了模型的对齐性能和奖励建模能力,相较基线有明显优势。

原文摘要 · Abstract (English)

Aligning Large Language Models (LLMs) with human preferences is crucial for their deployment in real-world applications. Recent advancements in Self-Rewarding Language Models suggest that an LLM can use its internal reward models (such as LLM-as-a-Judge) \cite{yuanself} to generate preference data, improving alignment performance without costly human annotation. However, we find that different internal reward models within the same LLM often generate inconsistent preferences. This inconsistency raises concerns about the reliability of self-generated preference data, hinders overall alignment performance, and highlights the need for further research to ensure reliable and coherent alignment with human preferences. To address this limitation, we propose Self-Consistent Internal Rewards (SCIR), a novel framework designed to enhance consistency among internal reward models during training. In each training step, we collect preference predictions from multiple pre-defined internal reward models and enforce consistency and confidence through an inconsistency penalty mechanism, thereby improving the reliability of these internal reward models. We selectively use data with consistent predictions for preference optimization, ensuring the quality of the preference data. By employing self-consistent internal rewards, our method significantly improves the alignment performance and reward modeling capability of LLMs, outperforming baseline methods by a notable margin.

自对齐奖励模型一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。