让模型更抗偏见:动态调整人类反馈的理性参数。
Mitigating Cognitive Bias in RLHF by Altering Rationality

- 用大模型判断人类反馈是否受认知偏见影响,动态调节理性参数。
- 在强偏见数据集上微调,仍能训练出更理性的下游模型。
- 适合关注人类反馈质量、强化学习鲁棒性的研究者。
在基于人类反馈的强化学习(RLHF)中,人类对模型输出的偏好被用来训练奖励模型,该模型为响应分配标量奖励值。由于奖励通过成对比较推断,其学习依赖于潜在奖励差异与观察偏好之间的假设关系,通常采用玻尔兹曼形式化,其中理性参数 beta 决定偏好反映奖励差异的一致性。实践中,beta 通常被视为固定常数,代表假设的统一标注者可靠性。然而,真实人类判断受认知偏见影响,导致系统性偏离奖励一致行为,且具上下文依赖性。为此,本文将理性参数视为上下文和标注相关的变量,设计一种利用大模型作为评判者,在奖励学习过程中动态调整 beta 的方法,有效降低可能受偏见或不可靠判断影响的比较权重。实验证明,该方法即使在强偏见偏好数据集上微调,也能学习到更理性的下游模型。
原文摘要 · Abstract (English)
How can we make models robust to even imperfect human feedback? In reinforcement learning from human feedback (RLHF), human preferences over model outputs are used to train a reward model that assigns scalar values to responses. Because these rewards are inferred from pairwise comparisons, this learning depends on an assumed relationship between latent reward differences and observed preferences, typically modeled using a Boltzmann formulation in which a rationality parameter beta informs how consistently preferences reflect reward differences. In practice, beta is typically treated as a fixed constant that reflects assumed uniform annotator reliability. However, human feedback is not this simplistic in practice: real human judgments are shaped by cognitive biases, leading to systematic deviations from reward-consistent behavior that arise contextually. To address this, we treat rationality as context- and annotation-dependent. We design an approach to dynamically adjust the rationality parameter beta during reward learning using an LLM-as-judge to assess the likely presence of cognitive biases. This approach effectively downweights comparisons that are likely to reflect biased or unreliable judgments. Empirically, we show that this approach learns a more rational downstream model, even when finetuning on datasets with strongly biased preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。