对比评估会放大模型隐性偏见,需谨慎使用。
To Compare, or Not to Compare: On Methodological Practices in Evaluating Social Bias

- 统一框架对比孤立评估与对比设置的差异
- 对比场景下偏见显著增强,尤其在思维链推理时
- 模型越大越易产生对比性偏见,适合评测者参考
随着大语言模型在关键应用中的普及,对其社会偏见进行稳健评估至关重要。然而,当前研究普遍存在方法学碎片化问题,导致结论矛盾。这主要源于忽视基准评估的结构设计。为此,我们提出一个统一可控的框架,将异构基准标准化,系统对比孤立评估与强制选择对比场景。关键在于分离出思维链推理、中立备选方案等结构干扰因素的影响。多模型家族评估显示存在显著的范式差距:孤立评估抑制偏见激活,而对比设置则成为潜在歧视的强催化剂,主因是上下文描述不明确。令人担忧的是,思维链推理在对比场景下加剧社会偏见,即使提供中立选项或声称随机回答,这种系统性偏见仍持续存在。最后,我们证明这种对比偏见具有泛化性,且随模型规模正向增长。因此,我们给出重要方法论建议:研究者应使用对比设置来审计隐藏偏见,但实践者不可在模糊现实任务中安全依赖对比部署。
原文摘要 · Abstract (English)
As Large Language Models are increasingly deployed in critical applications, robustly evaluating their social biases is paramount. However, the current literature suffers from widespread methodological fragmentation, which yields contradictory conclusions. This stems largely from ignoring the structural framing of benchmark-level evaluations. To resolve this, we introduce a unified and controllable framework that standardizes heterogeneous benchmarks to systematically contrast isolated demographic assessments with forced-choice comparative settings. Crucially, this allows us to disentangle the confounding effects of Chain-of-Thought reasoning, neutral fallback options, and other structural artifacts in social bias evaluations. Our evaluation across multiple model families reveals a massive, systematic paradigm gap: while isolated assessments limit prejudice activation, comparative settings act as aggressive catalysts for latent discrimination, a shift primarily driven by underspecified contexts. Alarmingly, CoT reasoning exacerbates social biases under comparative settings, and this systemic bias persists as a deterministic prejudice even when models are provided neutral fallback options or claim to answer randomly. Finally, we demonstrate that this comparative prejudice is a generalized phenomenon that scales positively with model size. Ultimately, we offer a crucial methodological guideline: while researchers must leverage comparative settings to robustly audit hidden biases, practitioners cannot safely rely on comparative deployments in ambiguous real-world tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。