用两次重写修正偏差,精准分析奖励模型偏好的真实原因
RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals
- 通过大模型重写生成不完美反事实样本,评估属性对奖励的影响
- 实验表明该方法能有效降低重写偏差,准确测量奖励模型敏感度
- 适合研究大模型对回应质量、情感等属性的因果偏好
奖励模型广泛用于对齐或评估大语言模型,作为人类偏好的代理。但奖励模型是黑箱,难以明确其实际奖励的内容。本文提出基于重写的属性处理估计器(RATE),用于衡量奖励模型对回应的高层次属性(如情感、帮助性、复杂性)的敏感性。重要的是,RATE能测量属性对奖励的因果影响。它利用大模型重写回应生成不完美的反事实样本,以评估因果效应。关键挑战在于此类重写存在偏差,可能严重影响估计结果。RATE的核心思想是通过双重重写来校正这种不完美重写带来的偏差。我们证明了该方法的合理性,并通过实验证明其有效性。
原文摘要 · Abstract (English)
Reward models are widely used as proxies for human preferences when aligning or evaluating LLMs. However, reward models are black boxes, and it is often unclear what, exactly, they are actually rewarding. In this paper we develop Rewrite-based Attribute Treatment Estimator (RATE) as an effective method for measuring the sensitivity of a reward model to high-level attributes of responses, such as sentiment, helpfulness, or complexity. Importantly, RATE measures the causal effect of an attribute on the reward. RATE uses LLMs to rewrite responses to produce imperfect counterfactuals examples that can be used to measure causal effects. A key challenge is that these rewrites are imperfect in a manner that can induce substantial bias in the estimated sensitivity of the reward model to the attribute. The core idea of RATE is to adjust for this imperfect-rewrite effect by rewriting twice. We establish the validity of the RATE procedure and show empirically that it is an effective estimator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。