发现神经网络同时用两种方式编码信息,提升可解释性。
Feature Integration Spaces: Joint Training Reveals Dual Encoding in Neural Network Representations
- 设计联合训练架构,同时捕捉特征身份与整合关系
- 重建误差降低51.6%,特征整合路径自发形成
- 小规模非线性模块即有效,适合模型可解释性研究
当前稀疏自编码器(SAE)方法假设神经网络激活可通过线性叠加分解为稀疏、可解释的特征。尽管重建精度高,但始终无法消除多义性并出现病理行为错误。我们提出神经网络在相同基底中以两种互补空间编码信息:特征身份与特征整合。为验证这一双编码假说,我们开发了顺序与联合训练架构,同步捕捉身份与整合模式。联合训练实现41.3%重建性能提升和51.6%的KL散度误差减少。该架构自发形成双模态特征组织:低平方范数特征参与整合路径,其余直接贡献于残差。仅3%参数的小型非线性组件即实现16.5%独立性能提升,表明其能高效捕捉关键计算关系。干预实验采用2×2因子刺激设计,显示整合特征对实验操控具有选择性敏感性,并在模型输出中引发系统性行为效应,包括跨语义维度的显著统计交互作用。本工作为神经表示中的双编码、有意义的非线性特征交互提供了系统证据,并从后验分析转向集成计算设计,奠定了下一代SAE的基础。
原文摘要 · Abstract (English)
Current sparse autoencoder (SAE) approaches to neural network interpretability assume that activations can be decomposed through linear superposition into sparse, interpretable features. Despite high reconstruction fidelity, SAEs consistently fail to eliminate polysemanticity and exhibit pathological behavioral errors. We propose that neural networks encode information in two complementary spaces compressed into the same substrate: feature identity and feature integration. To test this dual encoding hypothesis, we develop sequential and joint-training architectures to capture identity and integration patterns simultaneously. Joint training achieves 41.3% reconstruction improvement and 51.6% reduction in KL divergence errors. This architecture spontaneously develops bimodal feature organization: low squared norm features contributing to integration pathways and the rest contributing directly to the residual. Small nonlinear components (3% of parameters) achieve 16.5% standalone improvements, demonstrating parameter-efficient capture of computational relationships crucial for behavior. Additionally, intervention experiments using 2x2 factorial stimulus designs demonstrated that integration features exhibit selective sensitivity to experimental manipulations and produce systematic behavioral effects on model outputs, including significant statistical interaction effects across semantic dimensions. This work provides systematic evidence for (1) dual encoding in neural representations, (2) meaningful nonlinearly encoded feature interactions, and (3) introduces an architectural paradigm shift from post-hoc feature analysis to integrated computational design, establishing foundations for next-generation SAEs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。