arXiv:2411.16502cs.LGcs.AI2024-11ICLR被引 10

用对比样本解释语言奖励模型的决策依据。

Interpreting Language Reward Models via Contrastive Explanations

  • 通过修改高阶评价属性生成对比样本,分析模型局部行为。
  • 定量验证了方法能生成高质量对比解释。
  • 适合研究者分析奖励模型对不同属性的敏感性。

奖励模型(RMs)是将大语言模型输出与人类价值观对齐的关键组件。它们通过预测和比较不同回复的奖励分数,近似人类偏好。然而,由于通常是在大语言模型基础上添加标量输出头,这些模型本质上是黑箱,其预测难以解释。更透明的奖励模型有助于提升对大语言模型对齐的信任度。本文提出使用对比解释来说明奖励模型所做的任意二元判断。具体而言,我们生成一组与原始对比相似的新对比,以刻画奖励模型的局部行为。这些扰动后的回复被设计为显式修改人工指定的高层评估属性,使对奖励模型行为的分析建立在这些属性之上。在定量实验中,我们验证了该方法生成高质量对比解释的有效性。随后,我们展示了该方法在定性分析奖励模型对各评估属性的全局敏感性方面的价值,并演示了如何自动提取代表性样例,用于解释和比较不同奖励模型的行为。我们认为该方法是一个灵活的奖励模型解释框架,为更可解释、更可信的大语言模型对齐提供了基础。

原文摘要 · Abstract (English)

Reward models (RMs) are a crucial component in the alignment of large language models' (LLMs) outputs with human values. RMs approximate human preferences over possible LLM responses to the same prompt by predicting and comparing reward scores. However, as they are typically modified versions of LLMs with scalar output heads, RMs are large black boxes whose predictions are not explainable. More transparent RMs would enable improved trust in the alignment of LLMs. In this work, we propose to use contrastive explanations to explain any binary response comparison made by an RM. Specifically, we generate a diverse set of new comparisons similar to the original one to characterise the RM's local behaviour. The perturbed responses forming the new comparisons are generated to explicitly modify manually specified high-level evaluation attributes, on which analyses of RM behaviour are grounded. In quantitative experiments, we validate the effectiveness of our method for finding high-quality contrastive explanations. We then showcase the qualitative usefulness of our method for investigating global sensitivity of RMs to each evaluation attribute, and demonstrate how representative examples can be automatically extracted to explain and compare behaviours of different RMs. We see our method as a flexible framework for RM explanation, providing a basis for more interpretable and trustworthy LLM alignment.

奖励模型可解释性大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。