arXiv:2605.28639cs.CLcs.AI2026-05

模型抑制违规内容时,内部仍保留其表征并影响输出。

The Attentional White Bear Effect in Transformer Language Models

论文配图:The Attentional White Bear Effect in Transformer Language Models
图 1 · 摘自论文原文
  • 通过表征探测发现,禁用词在隐藏层仍可被准确恢复。
  • 即使不直接生成违禁词,注意力机制仍受其影响。
  • 适用于研究模型对齐与安全机制的可靠性。

指令式抑制广泛用于防止语言模型生成违规内容,但尚不清楚这种抑制是减少内部表征还是仅抑制表达。我们通过表征探测、注意力分析和行为语义泄漏实验,在多个Transformer模型中进行研究。结果表明,在抑制条件下,违规概念在隐藏表示中仍高度可恢复,持续影响注意力路由,并显著塑造下游生成内容,尽管实现了词汇层面的规避。这些现象在不同池化策略、间接语义控制及多个模型家族中均存在。研究揭示了行为对齐与表征对齐之间的根本性差距。

原文摘要 · Abstract (English)

Instruction-based suppression is widely used to prevent language models from generating prohibited content, yet it remains unclear whether suppression reduces internal representation or merely suppresses expression. We investigate this question through representational probing, attention analysis, and behavioral semantic leakage experiments across multiple transformer models. We find that prohibited concepts remain highly recoverable from hidden representations under suppression, continue to influence attention routing, and measurably shape downstream generations despite successful lexical avoidance. These effects persist across pooling strategies, indirect semantic controls, and multiple model families. Our results expose a fundamental gap between behavioral and representational alignment.

语言模型安全对齐注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。