提出新方法检测激活操控中的概念特异性和对齐泄漏,揭示现有验证方式的局限性。
SteerCheck: Attribution Specificity and Alignment Leakage in Activation-Steering Audits
- 设计预注册归因审计,分离多种类型的影响声明
- 发现94%的定向操控效果与符号余弦高度相关,25.3%样本余弦超0.5
- 适用于模型可解释性研究者,尤其关注生成模型操控可靠性
激活操控虽能改变行为,但未必针对特定概念。本文提出SteerCheck,一种预注册的归因审计方法,用于区分非目标KL、均值、受保护尾部、极性、迁移和语义等不同类型的主张。对960次Qwen3-14B干预进行精确重放显示:各向同性方向仅占据狭窄近正交区域,而符号随机化的同构方向常仍保留显著目标对齐。在符号随机化家族中,效应与符号余弦强相关(ρ=.94);25.3%的采样值余弦超过0.5,且所有超出观测均值效应的样本余弦均高于0.80。这种对齐泄漏不直接否定条件随机化检验,但限制了对照组的辨别能力,促使报告交换性假设、构造诊断指标$A$及经验余弦分布。主要Qwen完整门仍为负,因受保护尾部在所有家族中均失败。独立数据上,连续边缘转移仅在Qwen中出现,准确率转移则在任何选定单元中均未发生。前瞻性注册语言控制在Qwen和DeepSeek中通过完整门,但通过的DeepSeek去毒对照排除了类别分离;所有名义通过结果均对Γ=1.10敏感。冻结三评者开放式生成评估支持DeepSeek的事实修正,但不支持Qwen;自动判别器校准失败(宏平均F1=.562),因此全零语义结果仅为描述性。SteerCheck使这些条件性与混合结论可审计。
原文摘要 · Abstract (English)
Activation steering can change behaviour without establishing that the effect is specific to the intended concept. We introduce SteerCheck, a preregistered attribution audit that matches off-target KL and separates mean, protected-tail, polarity, transfer, and semantic claims. Exact replay of 960 Qwen3-14B interventions reveals complementary limits of common controls: isotropic directions occupy a narrow near-orthogonal region, whereas sign-randomized same-construction directions often retain substantial target alignment. Effect is strongly associated with signed cosine within the sign-randomized family ($ρ=.94$); $25.3\%$ of its draws exceed cosine $.5$, and every draw exceeding the observed mean effect has cosine above $.80$. This alignment leakage does not by itself invalidate a conditional randomization test; it limits what the comparator can distinguish and motivates reporting exchangeability assumptions, a construction diagnostic $A$, and the empirical cosine distribution. The primary Qwen complete gate remains negative because the protected tail fails all families. On independent data, continuous margin transfers only in Qwen and accuracy transfers in no selected cell. Prospectively registered language controls pass the complete gate in Qwen and DeepSeek, while a passing DeepSeek detox comparator rules out categorical separation; all nominal passes are sensitive to $Γ=1.10$. Frozen three-rater open-generation evaluation supports factual correction in DeepSeek but not Qwen; the automatic judge fails calibration (macro-F1 $.562$), so null-wide semantic results remain descriptive. SteerCheck makes these conditional and mixed conclusions auditable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。