测试大模型在不同语境下的偏见表现,发现同一模型在不同场景下偏见程度差异显著。
Contextual StereoSet: Stress-Testing Bias Alignment Robustness in Large Language Models
- 设计上下文可控的StereoSet基准,固定刻板印象内容但改变场景描述
- 1990年与2030年时间对比使所有模型偏见选择率上升(p<0.05)
- 适用于模型偏见评估的诊断与生产级筛查,强调条件敏感性
一个在实验室基准中避免刻板印象的模型,在实际部署中可能仍会表现出偏见。我们发现,当提示涉及不同地点、时间或受众时,偏见水平会发生显著变化——无需对抗性提示。为此,我们提出Contextual StereoSet:一个保持刻板印象内容不变,但系统性改变上下文框架的基准。在两个协议下测试13个模型,发现明显模式:锚定1990年(而非2030年)会使所有模型的刻板印象选择率显著上升(p<0.05);八卦式表述使6个全网格模型中有5个上升;外群体观察者视角可导致偏见变化达13个百分点。这些效应在招聘、贷款和求助情景中均复现。我们提出上下文敏感指纹(CSF):一种基于每维度离散度和成对对比的紧凑分析工具,包含自助法置信区间与多重检验校正。提供两种评估路径——涵盖360种情境的诊断网格,以及覆盖4,229个条目的预算协议,适用于深度分析与生产筛查。核心启示是方法论层面的:固定条件下的偏见评分未必具有泛化能力。这不是关于真实偏见率的断言,而是对评估鲁棒性的压力测试。CSF迫使评估者思考‘偏见在什么条件下出现?’,而非‘这个模型是否偏见?’。我们公开了基准、代码与结果。
原文摘要 · Abstract (English)
A model that avoids stereotypes in a lab benchmark may not avoid them in deployment. We show that measured bias shifts dramatically when prompts mention different places, times, or audiences -- no adversarial prompting required. We introduce Contextual StereoSet, a benchmark that holds stereotype content fixed while systematically varying contextual framing. Testing 13 models across two protocols, we find striking patterns: anchoring to 1990 (vs. 2030) raises stereotype selection in all models tested on this contrast (p<0.05); gossip framing raises it in 5 of 6 full-grid models; out-group observer framing shifts it by up to 13 percentage points. These effects replicate in hiring, lending, and help-seeking vignettes. We propose Context Sensitivity Fingerprints (CSF): a compact profile of per-dimension dispersion and paired contrasts with bootstrap CIs and FDR correction. Two evaluation tracks support different use cases -- a 360-context diagnostic grid for deep analysis and a budgeted protocol covering 4,229 items for production screening. The implication is methodological: bias scores from fixed-condition tests may not generalize.This is not a claim about ground-truth bias rates; it is a stress test of evaluation robustness. CSF forces evaluators to ask, "Under what conditions does bias appear?" rather than "Is this model biased?" We release our benchmark, code, and results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。