激活操控实际控制的是提取索引而非语义标签。
What Does Activation Steering Control? Attribution Across Answer Encodings and Output-Sensitive Subspaces

- 提出跨编码评估法,验证干预效果是否依赖答案编码方式。
- 在NormBank上,操控主要影响提取索引,而非语义标签或行位置。
- 仅15.4%的低秩输出敏感成分即可保留96.3%的操控效果,适合模型解释研究。
激活操控常在构建方向时使用的答案编码下进行评估,其报告的提升可能源于预期判断或与训练中看到的答案标识符的兼容性。我们提出跨编码操控评估(Cross-Encoding Steering Evaluation),在冻结干预方向的同时,对同一保留样本重新编码答案。在NormBank数据集上,当A/B/C标识符被重分配后,对比激活添加(CAA)在新映射下对提取索引的得分变化大于对语义标签的变化,这种现象称为提取索引跟随。改变标识符词汇(A/B/C、X/Y/Z、1/2/3)和行顺序发现,该效应随提取索引变化而非行位置。在层间方向范数匹配后,提取索引跟随主要出现在深层。一个包含方向平方范数15.4%的低秩输出敏感组件保留了96.3%的该效应。一种类推理时干预(ITI)方法在三个模型中也更倾向于提取索引而非语义标签跟随。总体而言,MNLI偏好提取索引跟随,而社交化学101(SC101)偏好语义标签跟随。多项选择与开放问答评估可能导致不同行为结论。因此,单一编码下的操控增益无法独立确定干预所控制的内容。
原文摘要 · Abstract (English)
Activation steering is often evaluated under the answer encoding used to construct the direction. A reported gain may reflect the intended judgment or compatibility with answer identifiers seen during construction. We introduce Cross-Encoding Steering Evaluation, which freezes an intervention while re-encoding answers to the same held-out items. On NormBank, after A/B/C identifiers are reassigned, contrastive activation addition (CAA) induces larger target-versus-source score changes for the extraction indices than for the semantic labels under the new mapping. We call this extraction-index following. Varying identifier vocabulary (A/B/C, X/Y/Z, or 1/2/3) and row order shows that the effect tracks extraction index rather than row position. After matching direction norms across layers, extraction-index following emerges mainly at later depths. A low-rank output-sensitive component containing 15.4% of the direction's squared norm retains 96.3% of this effect. An Inference-Time Intervention (ITI)-style method also favors extraction-index over semantic-label following on NormBank in three models. In aggregate, MNLI favors extraction-index following, whereas Social Chemistry 101 (SC101) favors semantic-label following. Multiple-choice and open-ended evaluations can yield different behavioral conclusions. Thus, a steering gain under one answer encoding does not by itself identify what the intervention controls.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。