arXiv:2606.18656cs.CL2026-06

发现大模型对刻板印象的过度安全响应会错误否定明确证据。

The Wrong Kind of Right: Quantifying and Localizing Misfired Alignment in LLMs

论文配图:The Wrong Kind of Right: Quantifying and Localizing Misfired Alignment in LLMs
图 1 · 摘自论文原文
  • 构建新基准VETO和MAR指标量化模型误判率
  • 25个模型均出现4.7%-18.9%的误触发对齐失败
  • 适合关注模型安全机制缺陷的研究者

本文研究大语言模型(LLMs)在对齐过程中产生的刻板印象与偏见问题,包含潜在敏感示例仅用于说明。研究表明,当前对齐机制可能导致模型错误拒绝明显由上下文支持的结论,即‘误触发对齐’。为此,我们提出VETO基准,包含2,032个基于BBQ的对比样本对,并定义新的衡量指标Misfired Alignment Rate(MAR),在0到100之间量化模型在刻板印象相关问题上的失败频率。对25个主流模型的测试显示,所有模型均表现出非零的MAR值(4.7%至18.9%),而人类参与者全部为0.0%。控制性提示实验表明,安全导向的提示可显著放大模型的误判率。机制分析发现,开放权重模型在后期层抑制了有依据的答案,且指令微调后的模型更易出现此类抑制。结果表明,当前对齐方法可能过度泛化表面安全信号,甚至压制客观证据,亟需更注重上下文一致性的对齐目标。

原文摘要 · Abstract (English)

Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment. Instead, this paper highlights the need for principled approaches to more advanced alignment. Alignment aims to ensure that large language models (LLMs) behave safely and reliably, including by avoiding unsafe inferences. However, we show that such safety-oriented behaviors can misfire: models may reject warranted conclusions even when they are explicitly supported by context. We call this failure mode misfired alignment, where alignment-induced changes cause LLMs to override explicit evidence. To quantify this phenomenon, specifically on stereotype-related alignment, we introduce VETO, a benchmark consisting of 2,032 BBQ-derived contrastive pairs, and define a new metric, Misfired Alignment Rate (MAR), which measures on a 0 to 100 scale how often a model fails on a stereotype-related question but succeeds on its contrastive counterpart. We benchmark 25 LLMs on VETO, and show that all LLMs, including the most recent ones, exhibit non-trivial (4.7 to 18.9%) MARs while all human participants achieve 0.0% MAR. Controlled priming experiments further show that alignment-induced cues can substantially amplify MAR across LLMs, indicating that these failures are not merely artifacts of individual examples but can be induced by safety-related framing. Mechanistic analyses on open-weight LLMs reveal late-layer suppression of evidence-supported answers, and comparisons between instruct and base LLMs suggest that this suppression emerges after instruction training. These findings show that current alignment methods can overgeneralize surface-level safety cues, to the point of overriding objective evidence, motivating more work on alignment objectives that better preserve contextual grounding.

大模型对齐偏见检测评估基准刻板印象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。