arXiv:2606.10931cs.CL2026-06

仅用一个偏见样本,就能让大模型系统性产生歧视性推理。

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO

论文配图:It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO
图 1 · 摘自论文原文
  • 通过单次GRPO训练,仅需一个偏见示例即可触发系统性偏差。
  • 模型对偏见的敏感度与其初始输出偏见概率相关。
  • 揭示了后训练对齐机制的致命弱点,适合安全与伦理研究者关注。

现代大型语言模型通常通过大规模后训练实现对齐,以确保行为公平可靠。本文研究了群体相对策略优化(GRPO)如何轻易突破这些防护机制。结果显示,仅需在单一偏见示例上进行一次GRPO训练,即可引发系统性偏见,且基于刻板印象的推理会泛化至不同属性、类别和评估基准。我们还发现,模型对偏见的易感性与其初始生成偏见输出的概率有关。结果表明,后训练对齐可被单个示例覆盖,暴露了其关键脆弱性。

原文摘要 · Abstract (English)

Warning: This paper contains several toxic and offensive statements. Modern large language models (LLMs) are typically aligned through large-scale post-training to ensure fair and reliable behavior. In this work, we investigate how easily such guardrails can be broken by Group Relative Policy Optimization (GRPO). We show that one-shot GRPO training on a single biased example is sufficient to induce systematic bias, with stereotype-driven reasoning generalizing across attributes, categories, and benchmarks. We further find that models differ in their susceptibility based on the initial likelihood of producing biased outputs. Our results reveal a critical vulnerability in post-training: alignment can be overridden by a single example.

模型对齐偏见传播安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。