构建首个系统化评估稀疏自编码器的基准,揭示现有指标与实际性能的脱节
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
- 设计8项跨领域评估指标,涵盖可解释性、特征解耦等维度
- 发现代理指标提升不等于实际性能改善,大模型规模下解耦优势更明显
- 开源200+个SAE模型,支持交互式多维对比分析
稀疏自编码器(SAEs)是解释语言模型激活的重要方法,近期研究大量聚焦于提升其有效性。然而,多数工作依赖缺乏实际意义的无监督代理指标进行评估。本文提出SAEBench,一个包含8项多样化指标的综合性评估框架,覆盖可解释性、特征解耦及去学习等实用场景。为实现系统性比较,我们开源了8种最新SAE架构和训练算法下的200多个SAE模型。评估表明,代理指标上的提升并不可靠地转化为实际性能优势。例如,马特里什卡SAE在现有代理指标上略逊一筹,但在特征解耦指标上显著优于其他架构,且该优势随模型规模增大而增强。SAEBench提供标准化评估框架,助力研究者分析缩放规律并细致比较不同SAE架构与训练方法。交互式界面支持在数百个开源SAE中灵活可视化指标间关系:www.neuronpedia.org/sae-bench
原文摘要 · Abstract (English)
Sparse autoencoders (SAEs) are a popular technique for interpreting language model activations, and there is extensive recent work on improving SAE effectiveness. However, most prior work evaluates progress using unsupervised proxy metrics with unclear practical relevance. We introduce SAEBench, a comprehensive evaluation suite that measures SAE performance across eight diverse metrics, spanning interpretability, feature disentanglement and practical applications like unlearning. To enable systematic comparison, we open-source a suite of over 200 SAEs across eight recently proposed SAE architectures and training algorithms. Our evaluation reveals that gains on proxy metrics do not reliably translate to better practical performance. For instance, while Matryoshka SAEs slightly underperform on existing proxy metrics, they substantially outperform other architectures on feature disentanglement metrics; moreover, this advantage grows with SAE scale. By providing a standardized framework for measuring progress in SAE development, SAEBench enables researchers to study scaling trends and make nuanced comparisons between different SAE architectures and training methodologies. Our interactive interface enables researchers to flexibly visualize relationships between metrics across hundreds of open-source SAEs at: www.neuronpedia.org/sae-bench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。