通过提取神经网络中的重复模式块,揭示模型对输入数据的结构化理解。
Discovering Chunks in Neural Embeddings for Interpretability
- 从神经网络隐状态中挖掘可识别的重复模式块,形成解释性字典。
- 在LLaMA等大模型中发现对应概念的嵌入状态,扰动可激活或抑制特定概念。
- 为理解复杂神经网络提供类人认知的解释框架,适合模型可解释性研究者。
理解神经网络因高维、相互作用的组件而困难。受人类认知中将复杂感官数据分块为重复实体的启发,我们提出利用该机制解释人工神经群体活动。生物与人工智能均需从结构化自然数据中学习,我们假设分块认知机制能为人工系统提供洞见。首先在训练于具规律性人工序列的循环神经网络(RNN)中验证,其隐状态反映这些模式,可提取为影响网络响应的块字典。扩展至大语言模型(如LLaMA),我们识别出对应输入概念的重复嵌入状态,对这些状态的扰动可激活或抑制相关概念。通过探索跨不同复杂度神经嵌入中提取可识别块字典的方法,我们的研究引入了一种新框架,将神经网络的群体活动视为其所处理数据的结构化映射。
原文摘要 · Abstract (English)
Understanding neural networks is challenging due to their high-dimensional, interacting components. Inspired by human cognition, which processes complex sensory data by chunking it into recurring entities, we propose leveraging this principle to interpret artificial neural population activities. Biological and artificial intelligence share the challenge of learning from structured, naturalistic data, and we hypothesize that the cognitive mechanism of chunking can provide insights into artificial systems. We first demonstrate this concept in recurrent neural networks (RNNs) trained on artificial sequences with imposed regularities, observing that their hidden states reflect these patterns, which can be extracted as a dictionary of chunks that influence network responses. Extending this to large language models (LLMs) like LLaMA, we identify similar recurring embedding states corresponding to concepts in the input, with perturbations to these states activating or inhibiting the associated concepts. By exploring methods to extract dictionaries of identifiable chunks across neural embeddings of varying complexity, our findings introduce a new framework for interpreting neural networks, framing their population activity as structured reflections of the data they process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。