用多智能体框架解析语言模型中一个神经元为何响应多个语义
NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models
- 将神经元解释转为迭代反馈过程,逐次拆解激活信号
- 相比单次分析方法,解释与实际激活相关性显著提升
- 适合研究大模型内部工作机制的开发者和研究人员
大型语言模型中的神经元级解释面临普遍存在的多义性挑战,即单个神经元对多个不同语义概念有响应。现有单次解释方法难以真实捕捉这种多概念行为。本文提出 NeuronScope,一种多智能体框架,将神经元解释重构为迭代、激活引导的过程。NeuronScope 显式将神经元激活分解为原子语义成分,聚类为不同语义模式,并通过神经元激活反馈迭代优化每项解释。实验表明,NeuronScope 能揭示隐藏的多义性,且生成的解释与实际激活的相关性显著高于单次基线方法。
原文摘要 · Abstract (English)
Neuron-level interpretation in large language models (LLMs) is fundamentally challenged by widespread polysemanticity, where individual neurons respond to multiple distinct semantic concepts. Existing single-pass interpretation methods struggle to faithfully capture such multi-concept behavior. In this work, we propose NeuronScope, a multi-agent framework that reformulates neuron interpretation as an iterative, activation-guided process. NeuronScope explicitly deconstructs neuron activations into atomic semantic components, clusters them into distinct semantic modes, and iteratively refines each explanation using neuron activation feedback. Experiments demonstrate that NeuronScope uncovers hidden polysemanticity and produces explanations with significantly higher activation correlation compared to single-pass baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。