用真实特征结构的合成数据测试稀疏自编码器,精准定位模型缺陷。
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
- 构建含相关性、层级与叠加特性的大规模合成数据集
- 发现匹配追踪型SAE利用噪声提升重建却未学真实特征
- 适合研究SAE架构缺陷与可解释性改进的学者
提升稀疏自编码器(SAEs)需要能精确验证架构创新的基准。现有基于大语言模型(LLM)的SAE基准噪声过大,难以区分改进效果;常见合成数据实验规模小、不标准化且不够真实。我们提出SynthSAEBench,一个用于评估SAEs的大规模合成数据基准与工具包,其数据具备真实特征特性:相关性、层级结构与叠加现象,并提供真实特征与激活的标注。该基准作为受控的下界测试:若SAE在构造上满足线性表示假设时仍失败,则在真实LLM上也难成功。该基准重现了已知的LLM SAE现象,包括重构与潜在表示质量间的脱节、探针性能差、以及由L0控制的精确率-召回率权衡。我们还发现一种新故障模式:匹配追踪型SAE利用叠加噪声提升重构性能,却不学习真实特征,表明更表达性强的编码方式易过拟合。SynthSAEBench为诊断SAE失效模式提供带真实标签的可控消融手段,并为架构设计提供明确目标。
原文摘要 · Abstract (English)
Improving Sparse Autoencoders (SAEs) requires benchmarks that can precisely validate architectural innovations. Current LLM-based SAE benchmarks are too noisy to differentiate architectural improvements, while commonly used synthetic-data experiments are too small-scale, unstandardized, and unrealistic to be meaningful. We introduce SynthSAEBench, a benchmark and toolkit for evaluating SAEs against large-scale synthetic data with realistic feature characteristics including correlation, hierarchy, and superposition, while providing ground-truth features and firings. SynthSAEBench acts as a controlled lower-bound test: SAE architectures that fail when the Linear Representation Hypothesis holds by construction have little hope on real LLMs. The benchmark reproduces known LLM SAE phenomena including the disconnect between reconstruction and latent quality, poor SAE probing, and a precision-recall trade-off mediated by L0, demonstrating that SynthSAEBench findings reproduce results on LLM SAEs. We further identify a novel failure mode: Matching Pursuit SAEs exploit superposition noise to improve reconstruction without learning ground-truth features, suggesting more expressive encoding procedures can easily overfit. SynthSAEBench complements LLM benchmarks with ground-truth features and controlled ablations for diagnosing SAE failure modes, while providing a clear target for SAE architecture work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。