arXiv:2510.08855cs.LGcs.AI2025-10中稿 · but the workshop d…被引 1

解决大模型稀疏自编码器训练中的特征吸收问题

Time-Aware Feature Selection: Adaptive Temporal Masking for Stable Sparse Autoencoder Training

  • 动态追踪激活强度、频率和重建贡献,生成随时间演化的特征重要性分数
  • 实验显示在Gemma-2-2b上吸收率显著低于TopK和JumpReLU方法
  • 适合需要稳定可解释特征的模型可靠性分析场景

理解大语言模型的内部表征对确保其可靠性与安全性至关重要,稀疏自编码器(SAEs)作为有前景的可解释性方法正受到关注。然而,现有SAE训练方法存在特征吸收问题,即特征(或神经元)相互吸收以最小化$L_1$惩罚,导致难以一致识别和分析模型行为。本文提出自适应时间掩码(ATM),通过追踪激活幅度、频率和重建贡献来计算随时间演化的特征重要性分数,并基于统计阈值实施概率掩码机制,实现更自然的特征选择。在Gemma-2-2b模型上的大量实验表明,ATM在保持优异重构质量的同时,吸收得分显著低于TopK和JumpReLU SAE等现有方法。该结果确立了ATM作为学习神经网络中稳定、可解释特征的原理性解决方案,为更可靠的模型分析提供了基础。

原文摘要 · Abstract (English)

Understanding the internal representations of large language models is crucial for ensuring their reliability and safety, with sparse autoencoders (SAEs) emerging as a promising interpretability approach. However, current SAE training methods face feature absorption, where features (or neurons) are absorbed into each other to minimize $L_1$ penalty, making it difficult to consistently identify and analyze model behaviors. We introduce Adaptive Temporal Masking (ATM), a novel training approach that dynamically adjusts feature selection by tracking activation magnitudes, frequencies, and reconstruction contributions to compute importance scores that evolve over time. ATM applies a probabilistic masking mechanism based on statistical thresholding of these importance scores, creating a more natural feature selection process. Through extensive experiments on the Gemma-2-2b model, we demonstrate that ATM achieves substantially lower absorption scores compared to existing methods like TopK and JumpReLU SAEs, while maintaining excellent reconstruction quality. These results establish ATM as a principled solution for learning stable, interpretable features in neural networks, providing a foundation for more reliable model analysis.

稀疏自编码器模型可解释性特征吸收大模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。