提升自动可解释性评估效率,让人工评测更准更便宜
Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability
- 用模型引导采样选关键输入,减少所需样本量13倍
- 用贝叶斯聚合降噪,每样本需评分减少3倍
- 整体评估成本降低40倍,支持大规模对比实验
解析神经元或激活空间方向是机制可解释性的重要课题。尽管已有多种自动化可解释方法提出,但其解释的可靠性仍不明确,且缺乏对方法准确性的客观评估。现有众包评估流程存在噪声大、成本高、仅关注最高激活输入等问题。本文提出两种新方法:第一,模型引导重要性采样(MG-IS),通过筛选最具信息量的输入,使达到同等评估精度所需的样本数减少约13倍;第二,贝叶斯评分聚合(BRAgg),有效缓解众包评分中的标签噪声,使每输入所需评分数减少约3倍。二者结合使评估成本降低约40倍,实现大规模评估可行性。最终,我们基于该框架对视觉网络中近期自动化可解释方法进行了大规模众包比较研究。
原文摘要 · Abstract (English)
Interpreting individual neurons or directions in activation space is an important topic in mechanistic interpretability. Numerous automated interpretability methods have been proposed to generate such explanations, but it remains unclear how reliable these explanations are, and which methods produce the most accurate descriptions. While crowd-sourced evaluations are commonly used, existing pipelines are noisy, costly, and typically assess only the highest-activating inputs, leading to unreliable results. In this paper, we introduce two techniques to enable cost-effective and accurate crowdsourced evaluation of automated interpretability methods beyond top activating inputs. First, we propose Model-Guided Importance Sampling (MG-IS) to select the most informative inputs to show human raters. In our experiments, we show this reduces the number of inputs needed to reach the same evaluation accuracy by ~13x. Second, we address label noise in crowd-sourced ratings through Bayesian Rating Aggregation (BRAgg), which allows us to reduce the number of ratings per input required to overcome noise by ~3x. Together, these techniques reduce the evaluation cost by ~40x, making large-scale evaluation feasible. Finally, we use our methods to conduct a large scale crowd-sourced study comparing recent automated interpretability methods for vision networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。