用形式语言测试文本自编码器的可解释性,发现因果特征需主动激励。
Analyzing (In)Abilities of SAEs via Formal Languages
- 在形式语言上训练稀疏自编码器,观察潜在空间的可解释特征
- 模型性能对训练偏差高度敏感,且相关特征未必有因果影响
- 提出应从底层激励学习因果特征,为未来研究提供方向
自编码器被广泛用于图像与文本领域中发现神经网络表征的可解释、解耦特征。尽管视觉领域的有效性与局限性已有充分研究,但文本领域的定性与定量结果仍不足。为此,我们在一个由形式语言构成的合成测试集(Dyck-2、Expr、English PCFG)上训练稀疏自编码器(SAEs),探索其在不同超参数设置下学习到的特征。结果显示,许多可解释的潜在特征会自然涌现。然而,与视觉领域类似,模型表现对训练流程的归纳偏置极为敏感。更重要的是,输入特征的相关潜在表示并不总能对模型计算产生因果影响。因此我们主张:因果性必须成为自编码器训练的核心目标,应从训练初期就激励学习因果相关特征。基于此,我们提出了初步方法,并在形式语言设定下进行了验证。
原文摘要 · Abstract (English)
Autoencoders have been used for finding interpretable and disentangled features underlying neural network representations in both image and text domains. While the efficacy and pitfalls of such methods are well-studied in vision, there is a lack of corresponding results, both qualitative and quantitative, for the text domain. We aim to address this gap by training sparse autoencoders (SAEs) on a synthetic testbed of formal languages. Specifically, we train SAEs on the hidden representations of models trained on formal languages (Dyck-2, Expr, and English PCFG) under a wide variety of hyperparameter settings, finding interpretable latents often emerge in the features learned by our SAEs. However, similar to vision, we find performance turns out to be highly sensitive to inductive biases of the training pipeline. Moreover, we show latents correlating to certain features of the input do not always induce a causal impact on model's computation. We thus argue that causality has to become a central target in SAE training: learning of causal features should be incentivized from the ground-up. Motivated by this, we propose and perform preliminary investigations for an approach that promotes learning of causally relevant features in our formal language setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。