arXiv:2606.02609cs.LGcs.AI2026-06

改进激活探针训练方法,提升可解释性模型的可靠性与易用性。

Building Better Activation Oracles

论文配图:Building Better Activation Oracles
图 1 · 摘自论文原文
  • 通过在线策略训练和优化对话数据集提升探针性能
  • 引入多层输入与改进注入公式,减少误判与模糊输出
  • 开源首个全面评估探针质量的基准工具AObench,适合可解释性研究者使用

激活探针(Activation Oracles, AOs)是解读残差流激活的有效方法,但现有方法存在幻觉和表述模糊等问题,且文本反转干扰使评估困难。为此,本文从四个方面改进AO训练流程:采用在线策略回放训练、优化对话数据集、引入更多网络层输入、改进注入公式。虽然性能提升有限,但使用体验显著改善。此外,我们开源了首个全面评估AO质量的基准工具AObench。整体上,本工作为可扩展、端到端可解释性模型的发展奠定了基础。

原文摘要 · Abstract (English)

Activation Oracles (AOs) are promising methods for interpreting residual stream activations. However, current AOs face important issues, such as hallucinations and vagueness. Additionally, text-inversion confounds make them hard to evaluate. To this end, we improve the Activation Oracle (AO) training regime in four ways: training on on-policy rollouts, improving the conversational dataset, feeding more layers and an improvement to the injection formula. The capability improvements are marginal, but quality of life improvements are quite substantial. In addition, we open source the first comprehensive evaluation suite for AO quality, which we call AObench. Overall, we hope that our work sets a foundation that helps improve AOs and other models in the paradigm of scalable, end-to-end interpretability.

可解释性激活探针评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。