研究大模型在否定指令下的反向思维反弹现象
Don't Think of the White Bear: Ironic Negation in Transformer Models Under Cognitive Load
- 通过不同干扰内容测试模型对否定词的反应强度
- 语义干扰越强反弹越明显,重复能有效抑制反弹
- 发现模型内部存在放大禁忌词的注意力机制
否定指令如‘不要提到$X$’会引发认知上的讽刺性反弹,使$X$更易被激活。大型语言模型(LLMs)同样面临此挑战:抑制概念需先内部激活它,反而可能诱发反弹。本文通过两个实验探究该矛盾:(1) 负载与内容:在否定指令后,变化干扰文本(语义、句法、重复),测量反弹强度;(2) 极性分离:检验模型是否区分同一概念的中性与负面表述,并判断这种区分是否预测反弹持续性。结果表明,反弹在否定后立即出现,且随语义干扰或更长干扰而增强,而重复则有助于抑制反弹。极性分离越强,反弹越持久。结合电路追踪分析,发现稀疏的中间层注意力头会放大禁用标记,而早期层则抑制。为此,我们发布ReboundBench数据集,包含5000个系统化设计的否定提示,用于探测大模型中的反弹现象。
原文摘要 · Abstract (English)
Negation instructions such as 'do not mention $X$' can paradoxically increase the accessibility of $X$ in human thought, a phenomenon known as ironic rebound. Large language models (LLMs) face the same challenge: suppressing a concept requires internally activating it, which may prime rebound instead of avoidance. We investigated this tension with two experiments. \textbf{(1) Load \& content}: after a negation instruction, we vary distractor text (semantic, syntactic, repetition) and measure rebound strength. \textbf{(2) Polarity separation}: We test whether models distinguish neutral from negative framings of the same concept and whether this separation predicts rebound persistence. Results show that rebound consistently arises immediately after negation and intensifies with longer or semantic distractors, while repetition supports suppression. Stronger polarity separation correlates with more persistent rebound. Together, these findings, complemented by a circuit tracing analysis that identifies sparse middle-layer attention heads amplifying forbidden tokens while early layers suppress, link cognitive predictions of ironic rebound with mechanistic insights into long-context interference. To support future work, we release ReboundBench, a dataset of $5,000$ systematically varied negation prompts designed to probe rebound in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。