arXiv:2505.11576cs.LGcs.AI2025-05NeurIPS被引 1

用认知中的分块原理,从神经网络中提取可理解的概念单元。

Concept-Guided Interpretability via Neural Chunking

  • 基于人类分块认知,将高维神经活动分解为可解释的概念单元。
  • 三种方法适配不同场景:有标签用聚类,无标签用无监督发现。
  • 提取的单元具有因果作用,替换后模型行为可预测改变。

神经网络常被视为黑箱,难以理解其内部运作。本文提出‘映射假说’:神经网络的原始群体活动模式反映了训练数据中的规律性。基于此,我们利用人类认知中的‘分块’机制,将高维神经动态划分为反映潜在概念的可解释单元。提出了三种互补方法:离散序列分块(DSC)在低维空间学习实体字典;群体平均(PA)提取对应已知标签的重复实体;无监督分块发现(UCD)适用于无标签场景。实验表明,这些方法能有效提取不依赖模型架构的概念编码单元,涵盖具体(词)、抽象(词性标注)和结构(叙事框架)等类型。此外,被提取的分块具有因果作用——将其‘嫁接’到模型中可引发可控且可预测的行为变化。本工作为可解释性研究开辟新方向,结合认知原则与自然数据结构,逐步揭示复杂学习系统的隐藏计算逻辑。

原文摘要 · Abstract (English)

Neural networks are often described as black boxes, reflecting the significant challenge of understanding their internal workings and interactions. We propose a different perspective that challenges the prevailing view: rather than being inscrutable, neural networks exhibit patterns in their raw population activity that mirror regularities in the training data. We refer to this as the Reflection Hypothesis and provide evidence for this phenomenon in both simple recurrent neural networks (RNNs) and complex large language models (LLMs). Building on this insight, we propose to leverage our cognitive tendency of chunking to segment high-dimensional neural population dynamics into interpretable units that reflect underlying concepts. We propose three methods to extract recurring chunks on a neural population level, complementing each other based on label availability and neural data dimensionality. Discrete sequence chunking (DSC) learns a dictionary of entities in a lower-dimensional neural space; population averaging (PA) extracts recurring entities that correspond to known labels; and unsupervised chunk discovery (UCD) can be used when labels are absent. We demonstrate the effectiveness of these methods in extracting concept-encoding entities agnostic to model architectures. These concepts can be both concrete (words), abstract (POS tags), or structural (narrative schema). Additionally, we show that extracted chunks play a causal role in network behavior, as grafting them leads to controlled and predictable changes in the model's behavior. Our work points to a new direction for interpretability, one that harnesses both cognitive principles and the structure of naturalistic data to reveal the hidden computations of complex learning systems, gradually transforming them from black boxes into systems we can begin to understand.

可解释性神经分块认知机制概念提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。