让不同初始化的BERT模型提取出通用特征,提升可解释性。
Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders

- 用正交普鲁克斯特旋转对齐不同种子的激活空间,再联合训练稀疏自编码器。
- 跨种子特征相关性达0.70以上,优于传统后处理对齐方法。
- 适合研究模型内部机制、社会语言学模式等可解释性方向的学者。
我们提出一种基于普鲁克斯特约束的联合端到端Top-K稀疏自编码器(SAE),用于从独立训练的BERT模型中提取跨种子通用特征。在机制可解释性中,跨种子特征通用性是一个根本挑战:由于字典学习是非凸的,不同初始化的网络会学习到错位的特征空间,导致看似相同的特征实际存在随机初始化差异。我们通过在联合SAE训练前计算种子间激活空间的正交普鲁克斯特旋转来解决此问题,结合Top-K稀疏性、端到端下游优化以及基于先前SAE文献的死特征恢复辅助损失。在三个基准数据集(SST-2、Stanford Politeness、TweetEval Emotion)上评估五组独立种子对(共十台BERT模型),我们的完整流程在所有数据集上均产生更通用的特征(跨种子皮尔逊相关系数≥0.70)。初步定性分析确认,高通用性特征编码了可解释的社会语言学模式。
原文摘要 · Abstract (English)
We present a Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoder (SAE) for extracting cross-seed universal features from independently trained BERT models. Cross-seed feature universality is a fundamental challenge in mechanistic interpretability: because dictionary learning is non-convex, independently trained networks learn misaligned feature spaces, so apparently identical features may differ by random initialization. We address this by computing an orthogonal Procrustes rotation between seeds' activation spaces before joint SAE training, combining Top-K sparsity, end-to-end downstream optimization, and an auxiliary dead-feature revival loss based on previous SAE literature. Evaluating on five independent seed pairs (ten BERT models) across three benchmark datasets (SST-2, Stanford Politeness, TweetEval Emotion), our full pipeline produces more universal features (Pearson r $\geq$ 0.70 across seeds) than post-hoc alignment baselines on all three datasets. A minimal qualitative analysis confirms that high-universality features encode interpretable sociolinguistic patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。