自动评估大模型特征描述的准确性,揭示错误描述的根源。
FADE: Why Bad Descriptions Happen to Good Features
- 提出FADE框架,自动评估特征与描述的对齐程度。
- 发现SAE特征比MLP神经元更难生成准确描述。
- 适用于提升自动化可解释性工具的质量,适合研究者参考。
近期机制可解释性进展表明,自动化分析大模型隐空间表示的可解释性流程具有潜力。然而,该领域缺乏标准化方法来评估所发现特征的有效性。本文提出FADE:特征对齐描述评估框架,一种可扩展、模型无关的自动评估方法,用于衡量特征与描述之间的对齐度。FADE基于四个核心指标——清晰度、响应性、纯度和忠实性,系统量化特征与描述不一致的原因。我们将FADE应用于现有开源特征描述,并评估自动化可解释性流水线的关键组件,以提升描述质量。研究发现,生成特征描述面临根本性挑战,尤其在SAE相对于MLP神经元时更为明显,为自动化可解释性的局限与未来方向提供了洞见。FADE已开源,项目地址:https://github.com/brunibrun/FADE。
原文摘要 · Abstract (English)
Recent advances in mechanistic interpretability have highlighted the potential of automating interpretability pipelines in analyzing the latent representations within LLMs. While this may enhance our understanding of internal mechanisms, the field lacks standardized evaluation methods for assessing the validity of discovered features. We attempt to bridge this gap by introducing FADE: Feature Alignment to Description Evaluation, a scalable model-agnostic framework for automatically evaluating feature-to-description alignment. FADE evaluates alignment across four key metrics - Clarity, Responsiveness, Purity, and Faithfulness - and systematically quantifies the causes of the misalignment between features and their descriptions. We apply FADE to analyze existing open-source feature descriptions and assess key components of automated interpretability pipelines, aiming to enhance the quality of descriptions. Our findings highlight fundamental challenges in generating feature descriptions, particularly for SAEs compared to MLP neurons, providing insights into the limitations and future directions of automated interpretability. We release FADE as an open-source package at: https://github.com/brunibrun/FADE
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。