arXiv:2601.08070cs.AIcs.CL2026-01被引 1

负向指令常失效,因提禁词反而让模型更想用它。

Semantic Gravity Wells: Why Negative Constraints Backfire

  • 用语义压力量化模型生成禁词的内在倾向
  • 失败时抑制力仅成功时的1/4.4,差距显著
  • 提禁词会激活目标词,晚层网络还反向增强

负向约束(如‘不要使用词X’)是检验大模型指令遵循能力的关键挑战。尽管形式简单,但此类指令失败率极高,且失败机制尚不明确。本文首次系统性探究负向指令失效机理,提出语义压力作为衡量模型生成禁词概率的量化指标,发现违反概率与压力呈紧密逻辑关系(p=σ(-2.40+2.27·P₀);n=40,000样本;斜率95%置信区间[2.21,2.33])。通过逐层分析发现,失败时指令对目标词的压制力度仅为成功时的5.2个百分点,而成功时达22.8点,存在4.4倍差异。进一步揭示两种失效模式:在87.5%的案例中,指令提及禁词反而激活目标表示(引子失效);在12.5%中,晚期前馈网络贡献+0.39至目标概率,近为成功情况的4倍,压倒早期抑制信号。激活修补实验确认第23–27层为因果关键:替换其激活可反转约束效果。研究揭示负向指令设计的根本矛盾:命名禁词本身即引发模型生成冲动。

原文摘要 · Abstract (English)

Negative constraints (instructions of the form "do not use word X") represent a fundamental test of instruction-following capability in large language models. Despite their apparent simplicity, these constraints fail with striking regularity, and the conditions governing failure have remained poorly understood. This paper presents the first comprehensive mechanistic investigation of negative instruction failure. We introduce semantic pressure, a quantitative measure of the model's intrinsic probability of generating the forbidden token, and demonstrate that violation probability follows a tight logistic relationship with pressure ($p=σ(-2.40+2.27\cdot P_0)$; $n=40{,}000$ samples; bootstrap $95%$ CI for slope: $[2.21,,2.33]$). Through layer-wise analysis using the logit lens technique, we establish that the suppression signal induced by negative instructions is present but systematically weaker in failures: the instruction reduces target probability by only 5.2 percentage points in failures versus 22.8 points in successes -- a $4.4\times$ asymmetry. We trace this asymmetry to two mechanistically distinct failure modes. In priming failure (87.5% of violations), the instruction's explicit mention of the forbidden word paradoxically activates rather than suppresses the target representation. In override failure (12.5%), late-layer feed-forward networks generate contributions of $+0.39$ toward the target probability -- nearly $4\times$ larger than in successes -- overwhelming earlier suppression signals. Activation patching confirms that layers 23--27 are causally responsible: replacing these layers' activations flips the sign of constraint effects. These findings reveal a fundamental tension in negative constraint design: the very act of naming a forbidden word primes the model to produce it.

大模型指令遵循负向约束机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。