arXiv:2606.30989cs.CLcs.AI2026-06

模型推理时会误用群体统计偏见,导致对个体的不公平判断。

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG

论文配图:Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG
图 1 · 摘自论文原文
  • 发现模型在推理中将群体规律错误套用于个体,产生看似合理却偏见的结论。
  • 提出公平注入框架,用Fair-GCG自动发现有效纠偏短语,提升多场景公平性。
  • 效果可跨模型大小迁移,适用于开放生成与真实敏感任务。

尽管近期大语言模型的推理能力有所提升,但公平性问题仍存。本文识别出一种新故障模式——演绎式刻板印象:模型将群体层面的统计规律错误应用于个体案例,生成逻辑自洽却社会偏见严重的推论。我们从统计角度解释该现象,并提出一种推理时注入的纠正框架。进一步设计Fair-GCG系统,用于系统性发现有效的纠正短语。经Fair-GCG发现的注入短语,在多个公平性基准上均表现更优,具备从小模型到大模型的泛化能力,显著提升推理层级的公平性,减少开放式生成中的偏见,并成功迁移到真实世界中的公平敏感任务。

原文摘要 · Abstract (English)

Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language models (LLMs), failures persist. In this work, we identify a failure mode, deductive stereotyping, in which models apply population-level statistical regularities to individual cases, producing logically coherent yet socially biased inferences. We provide a statistical interpretation of this phenomenon. To steer models toward fairness-aware reasoning, we propose a reasoning-time injection framework. We further introduce Fair-GCG to systematically discover effective injection phrases. Injection phrases discovered by Fair-GCG improve performance across multiple fairness benchmarks, generalize from smaller to larger LLMs, improves reasoning-level fairness, reduces bias in open-ended generation, and transfer to real-world fairness-sensitive tasks.

公平性推理偏差大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。