用大模型自动发现奖励模型的隐藏偏见,提升AI对齐可靠性。
Automatically Finding Reward Model Biases
- 让大模型迭代生成并优化潜在偏见假设
- 在Skywork-V2-8B中发现冗余空格和幻觉内容被错误奖励
- 适合研究模型对齐与自动化可解释性的学者
奖励模型在大语言模型后训练中起核心作用。然而以往研究发现,它们可能奖励无关或不良特征,如文本长度、格式、幻觉和阿谀奉承。本文提出并研究了在自然语言中自动发现奖励模型偏见的问题。我们提出一种简单方法:使用大模型迭代生成并优化候选偏见。该方法能复现已知偏见,并发现新偏见,例如在领先的开源奖励模型Skywork-V2-8B中,发现其常错误偏好包含冗余空格的回复及包含幻觉内容的回复。此外,我们证明进化式迭代优于一次性最佳N选一搜索,并通过人为注入偏见验证了管道的召回能力。希望本工作推动通过自动化可解释性方法改进奖励模型的研究。
原文摘要 · Abstract (English)
Reward models are central to large language model (LLM) post-training. However, past work has shown that they can reward spurious or undesirable attributes such as length, format, hallucinations, and sycophancy. In this work, we introduce and study the research problem of automatically finding reward model biases in natural language. We offer a simple approach of using an LLM to iteratively propose and refine candidate biases. Our method can recover known biases and surface novel ones: for example, we found that Skywork-V2-8B, a leading open-weight reward model, often mistakenly favors responses with redundant spacing and responses with hallucinated content. In addition, we show evidence that evolutionary iteration outperforms flat best-of-N search, and we validate the recall of our pipeline using synthetically injected biases. We hope our work contributes to further research on improving RMs through automated interpretability methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。