PRISM让神经网络能同时识别多个语义,更真实地描述大模型内部机制。
Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
- 为每个神经元生成多语义描述,突破单一概念假设
- 在基准测试中显著提升描述准确性和多义性捕捉能力
- 适合研究大模型内部表征的学者与可解释性开发者
自动化可解释性研究旨在识别神经网络特征中编码的概念,以增进对模型行为的理解。在自然语言处理的大语言模型(LLM)背景下,现有的神经元级特征描述方法面临两大挑战:鲁棒性不足,以及默认每个神经元只编码单一概念(单义性),而越来越多证据表明存在多义性。这一假设限制了特征描述的表达力,难以全面捕捉模型内部的行为。为此,我们提出多概念特征识别与评分方法(PRISM),一种专为捕捉大语言模型特征复杂性设计的新框架。不同于多数NLP自动化可解释性方法中每个神经元仅分配一个描述的做法,PRISM能够生成更细致的描述,兼顾单义与多义行为。我们将PRISM应用于大语言模型,并通过广泛基准测试证明,该方法生成的特征描述更准确、更忠实,整体描述质量(描述得分)和多义性情境下的概念区分能力(多义性得分)均显著提升。
原文摘要 · Abstract (English)
Automated interpretability research aims to identify concepts encoded in neural network features to enhance human understanding of model behavior. Within the context of large language models (LLMs) for natural language processing (NLP), current automated neuron-level feature description methods face two key challenges: limited robustness and the assumption that each neuron encodes a single concept (monosemanticity), despite increasing evidence of polysemanticity. This assumption restricts the expressiveness of feature descriptions and limits their ability to capture the full range of behaviors encoded in model internals. To address this, we introduce Polysemantic FeatuRe Identification and Scoring Method (PRISM), a novel framework specifically designed to capture the complexity of features in LLMs. Unlike approaches that assign a single description per neuron, common in many automated interpretability methods in NLP, PRISM produces more nuanced descriptions that account for both monosemantic and polysemantic behavior. We apply PRISM to LLMs and, through extensive benchmarking against existing methods, demonstrate that our approach produces more accurate and faithful feature descriptions, improving both overall description quality (via a description score) and the ability to capture distinct concepts when polysemanticity is present (via a polysemanticity score).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。