arXiv:2602.01716cs.CL2026-02

用模型内部信号预测大模型行为引导是否有效

Mechanistic Indicators of Steering Effectiveness in Large Language Models

  • 通过熵和KL散度分析模型中间激活,判断引导效果
  • 发现有效引导对应熵结构保留与目标概念对齐
  • 为两种主流引导方法提供更强评估基准

基于激活值的引导技术可在不重新训练的情况下让大语言模型表现出特定行为。尽管应用广泛,但决定引导成功或失败的机制仍不明确,以往研究多依赖黑箱输出或由大模型判断。本研究探讨是否可利用模型内部信号诊断引导可靠性。聚焦两种信息论指标:基于熵的归一化分支因子(NBF)和引导后激活与词汇空间中目标概念间的KL散度。假设有效引导表现为解码过程中熵结构的保持与KL对齐的一致性。基于两套架构不同的大模型间评委一致性高的验证,我们以大模型生成的标注作为真实标签,证明这些机制信号能有效预测引导成功与否并估计失败概率。此外,我们为对比激活添加(CAA)和稀疏自编码器引导这两种最常用方法引入更强的评估基准。

原文摘要 · Abstract (English)

Activation-based steering enables Large Language Models (LLMs) to exhibit targeted behaviors by intervening on intermediate activations without retraining. Despite its widespread use, the mechanistic factors that govern when steering succeeds or fails remain poorly understood, as prior work has relied primarily on black-box outputs or LLM-based judges. In this study, we investigate whether the reliability of steering can be diagnosed using internal model signals. We focus on two information-theoretic measures: the entropy-derived Normalized Branching Factor (NBF), and the Kullback-Leibler (KL) divergence between steered activations and targeted concepts in the vocabulary space. We hypothesize that effective steering corresponds to structured entropy preservation and coherent KL alignment across decoding steps. Building on a reliability study demonstrating high inter-judge agreement between two architecturally distinct LLMs, we use LLM-generated annotations as ground truth and show that these mechanistic signals provide meaningful predictive power for identifying successful steering and estimating failure probability. We further introduce a stronger evaluation baseline for Contrastive Activation Addition (CAA) and Sparse Autoencoder-based steering, the two most widely adopted activation-steering methods.

大模型引导机制分析信息论评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。