揭示神经网络如何用线性叠加表示信息,并提取可解释的稀疏代码。
From superposition to sparse codes: interpretable representations in neural networks
- 基于可识别性理论,分类模型恢复潜在特征
- 稀疏编码从激活中提取解耦特征,符合压缩感知原理
- 定量指标验证提取特征与人类理解的一致性
理解神经网络中的信息表征是神经科学与人工智能的核心挑战。尽管具有非线性结构,近期研究表明神经网络以超位置方式编码特征,即输入概念在线性叠加于网络表征中。本文提出一个新视角,解释该现象并为从神经激活中提取可解释表征奠定基础。理论框架包含三步:(1) 可识别性理论表明,训练用于分类的神经网络能恢复潜在特征,至多为线性变换;(2) 稀疏编码方法可利用压缩感知原则,从这些表征中提取解耦特征;(3) 定量可解释性度量提供评估手段,确保提取特征与人类可理解概念一致。通过融合理论神经科学、表征学习与可解释性研究,本文提出理解人工与生物系统中神经表征的新兴视角,对神经编码理论、人工智能透明化及深度学习可解释性具有深远意义。
原文摘要 · Abstract (English)
Understanding how information is represented in neural networks is a fundamental challenge in both neuroscience and artificial intelligence. Despite their nonlinear architectures, recent evidence suggests that neural networks encode features in superposition, meaning that input concepts are linearly overlaid within the network's representations. We present a perspective that explains this phenomenon and provides a foundation for extracting interpretable representations from neural activations. Our theoretical framework consists of three steps: (1) Identifiability theory shows that neural networks trained for classification recover latent features up to a linear transformation. (2) Sparse coding methods can extract disentangled features from these representations by leveraging principles from compressed sensing. (3) Quantitative interpretability metrics provide a means to assess the success of these methods, ensuring that extracted features align with human-interpretable concepts. By bridging insights from theoretical neuroscience, representation learning, and interpretability research, we propose an emerging perspective on understanding neural representations in both artificial and biological systems. Our arguments have implications for neural coding theories, AI transparency, and the broader goal of making deep learning models more interpretable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。