自解释评分受评估流程影响极大,跨论文比较可能误判模型性能。
Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance

- 用多模型多方法实验发现评估流程方差大于模型架构方差
- 检测任务最稳定,模糊测试完全不可靠,平均分掩盖个体特征不稳
- 提出方差分解与稳定性检查工具,助研究者识别评估陷阱
稀疏自编码器(SAE)可解释性跨论文比较常依赖自解释评分:由语言模型生成特征解释,再由另一语言模型评分。为使比较有意义,评分需反映特征本身属性而非评估流程偏差。我们通过四类指标(模拟、检测、模糊测试、纯净度)、两个模型(Pythia-160M、Apertus-8B)及四方面方法变体系统实验发现,该假设不成立。具体显示:1)方法论方差在所有指标与模型下均超过架构方差;2)各指标稳定性各异,检测最稳,模糊测试全条件下不可靠;3)顶K特征排名在不同语料和采样条件下不一致,导致单个特征不稳定被平均分掩盖,仅监控解释相似性无法察觉此问题。结论表明,基于自解释评分的跨论文比较可能反映的是评估流程差异而非模型架构差异,对SAE效用争论有重要影响。为此,我们提供方差分解方法、稳定性检查工具及最低报告清单,以支持更可靠的可解释性评估。
原文摘要 · Abstract (English)
Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a language model (LM) explains each feature, and another LM scores the explanation. For these comparisons to be meaningful, scores must reflect stable properties of the features rather than confounding aspects of the evaluation pipeline. Through systematic experiments across four metrics (simulation, detection, fuzzing, purity), two models (Pythia-160M, Apertus-8B), and four axes of methodological variation, we show that this assumption does not hold. Specifically, we find that R1) methodological variance collectively exceeds architectural variance across all metrics and tested models; R2) each metric exhibits a distinct instability profile, with detection being the most stable and fuzzing unreliable across all conditions; R3) top-k feature rankings do not stay consistent across corpus and draw conditions, masking per-feature instability behind stable mean scores; a failure that cannot be detected by monitoring explanation similarity alone. These findings suggest that cross-paper comparisons based on autointerpretability scores may reflect pipeline differences rather than architectural differences, with implications for the ongoing debate on SAE utility. More broadly, unreliable evaluation slows progress in interpretability research at a time when reliable tools for understanding AI systems are needed. To support evaluation, we contribute a variance decomposition approach, a Stability Check, and a Minimum Reporting Checklist.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。