arXiv:2503.10894cs.CLcs.AI2025-03ICLR被引 10

用超网络自动定位模型中概念的实现位置并构建特征。

HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks

  • 基于Transformer的超网络自动搜索概念在隐藏状态中的位置。
  • 在Llama3-8B上达到RAVEL基准最佳性能。
  • 适合关注可解释性自动化与模型内部机制研究者。

机制可解释性已取得显著进展,能够识别神经网络中中介概念(如人物出生年份)的特征(如隐藏激活空间中的方向),并实现可预测操控。分布式对偶搜索(DAS)利用反事实数据监督,在隐藏状态中学习概念特征,但其假设可承受对潜在特征位置进行暴力搜索。为解决此问题,我们提出HyperDAS,一种基于Transformer的超网络架构,能(1)自动定位残差流中概念实现的标记位置,(2)为这些残差流向量构建概念特征。在Llama3-8B上的实验表明,HyperDAS在RAVEL基准上实现了最先进的解耦概念性能。此外,我们回顾了设计决策,以缓解HyperDAS(如同所有强大可解释方法)可能向目标模型注入新信息而非忠实解读的风险。

原文摘要 · Abstract (English)

Mechanistic interpretability has made great strides in identifying neural network features (e.g., directions in hidden activation space) that mediate concepts(e.g., the birth year of a person) and enable predictable manipulation. Distributed alignment search (DAS) leverages supervision from counterfactual data to learn concept features within hidden states, but DAS assumes we can afford to conduct a brute force search over potential feature locations. To address this, we present HyperDAS, a transformer-based hypernetwork architecture that (1) automatically locates the token-positions of the residual stream that a concept is realized in and (2) constructs features of those residual stream vectors for the concept. In experiments with Llama3-8B, HyperDAS achieves state-of-the-art performance on the RAVEL benchmark for disentangling concepts in hidden states. In addition, we review the design decisions we made to mitigate the concern that HyperDAS (like all powerful interpretabilty methods) might inject new information into the target model rather than faithfully interpreting it.

可解释性超网络机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。