arXiv:2605.24583cs.LGcs.CL2026-05

纠正对模型对齐效果的误判,避免模板差异带来的干扰。

Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol

  • 设计模板控制的差中差协议,分离对齐变化与对话格式影响。
  • 实测显示有效秩被高估2.0-3.9倍,差中差法恢复真实拒绝方向。
  • 适用于研究对齐训练如何改变模型内部激活的科研人员。

对比对齐前后模型在安全相关输入下的内部激活,是理解安全训练影响的自然方法。传统做法直接计算对齐模型与基线模型激活的差值矩阵,但该方法存在偏差:对齐模型在聊天模板下评估,而基线模型从未见过此模板,导致差值同时包含对齐变化与格式差异。本文提出四类分解方案(原始、模板控制、对齐内、差中差),可分离二者影响。仅模板控制即可消除Llama-3.1-8B、Gemma-2-9B和Qwen-2.5-7B上2.0–3.9倍的有效秩膨胀;差中差对比成功复现Arditi等(2024)的拒绝方向,余弦相似度从0.18–0.39提升至0.50–0.86。跨模型投影消融验证恢复子空间具有行为活性,且奇异值顺序非因果顺序。在可控测试床验证后,提炼出对激活差值研究的测量建议。

原文摘要 · Abstract (English)

Comparing a model's internal activations before and after alignment is a natural way to ask what safety training changes: one forms the matrix of paired aligned-minus-base activations on safety-relevant inputs and reads off its effective rank or top direction. We show the obvious way to form this matrix is confounded. The aligned model is evaluated under a chat template the base model never saw, so the naive difference conflates the alignment shift with chat formatting. We introduce a four-variant decomposition of the modification matrix (naive, template-controlled, within-aligned, and difference-in-differences, DiD) that separates the two effects. Template control alone removes a 2.0-3.9x inflation of the measured effective rank across Llama-3.1-8B, Gemma-2-9B, and Qwen-2.5-7B; the DiD contrast is what recovers the refusal direction of Arditi et al. (2024), lifting its cosine alignment from 0.18-0.39 to 0.50-0.86. Projection-ablation across the three families confirms the recovered subspace is behaviorally active and that singular-value order is not causal order. We validate the protocol on a controlled testbed and distill it into measurement recommendations for activation-difference studies of alignment.

模型对齐激活分析差中差法评测方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。