发现语言模型奖励模型对查询前缀的敏感偏差,影响公平性
Detecting Prefix Bias in LLM-based Reward Models
- 通过分析查询前缀微小变化对模型偏好影响,检测系统性偏差
- 在多个开源数据集上发现性别和种族维度存在显著偏好偏差
- 提出数据增强策略可有效降低前缀偏差,适合关注AI公平性的研究者
基于人类反馈的强化学习(RLHF)已成为利用人类偏好数据对语言模型进行任务特定微调的关键范式。尽管已有大量公开的偏好数据集提供响应间的成对比较,但由此产生的奖励模型中的潜在偏差仍缺乏深入探索。本文提出新方法,检测并评估基于大语言模型的奖励模型中由查询前缀微小变化引发的系统性偏好偏移(即前缀偏差)。我们利用这些度量工具,在多个开源偏好数据集与不同架构的奖励模型上揭示了显著的性别与种族维度偏差。实验表明,该类偏差在不同模型架构下普遍存在。此外,我们提出一种数据增强策略以缓解此类偏差,实验证明其能有效降低前缀偏差的影响。研究强调了在设计与评估奖励模型时需具备偏差意识,推动更公平可靠的AI系统发展。
原文摘要 · Abstract (English)
Reinforcement Learning with Human Feedback (RLHF) has emerged as a key paradigm for task-specific fine-tuning of language models using human preference data. While numerous publicly available preference datasets provide pairwise comparisons of responses, the potential for biases in the resulting reward models remains underexplored. In this work, we introduce novel methods to detect and evaluate prefix bias -- a systematic shift in model preferences triggered by minor variations in query prefixes -- in LLM-based reward models trained on such datasets. We leverage these metrics to reveal significant biases in preference models across racial and gender dimensions. Our comprehensive evaluation spans diverse open-source preference datasets and reward model architectures, demonstrating susceptibility to this kind of bias regardless of the underlying model architecture. Furthermore, we propose a data augmentation strategy to mitigate these biases, showing its effectiveness in reducing the impact of prefix bias. Our findings highlight the critical need for bias-aware dataset design and evaluation in developing fair and reliable reward models, contributing to the broader discourse on fairness in AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。