arXiv:2510.01048cs.CLcs.AI2025-10中稿 · The Eight Workshop…综述被引 5

用自然语言解释大模型内部组件,让黑箱决策更透明

Interpreting Language Models Through Concept Descriptions: A Survey

  • 用生成模型自动描述神经元等组件的语义概念
  • 现有方法能生成可读性强的概念描述,但评估标准不统一
  • 适合想提升模型可解释性的研究者和工程师

理解神经网络的决策过程是机制可解释性的重要目标。在大语言模型(LLMs)中,这涉及揭示底层机制,识别神经元、注意力头等组件的作用,以及稀疏自编码器(SAEs)提取的稀疏特征等抽象表示。近年来,大量工作通过强大生成模型,为这些组件生成开放词汇、自然语言的概念描述来应对这一挑战。本文首次系统综述了组件概念描述这一新兴领域,梳理了生成描述的关键方法、自动化与人工评估指标的演进,以及支撑该研究的数据集。我们的分析揭示了对更严格、因果性评估的迫切需求。通过总结现状并指出关键挑战,本综述为未来提升模型透明度的研究提供了路线图。

原文摘要 · Abstract (English)

Understanding the decision-making processes of neural networks is a central goal of mechanistic interpretability. In the context of Large Language Models (LLMs), this involves uncovering the underlying mechanisms and identifying the roles of individual model components such as neurons and attention heads, as well as model abstractions such as the learned sparse features extracted by Sparse Autoencoders (SAEs). A rapidly growing line of work tackles this challenge by using powerful generator models to produce open-vocabulary, natural language concept descriptions for these components. In this paper, we provide the first survey of the emerging field of concept descriptions for model components and abstractions. We chart the key methods for generating these descriptions, the evolving landscape of automated and human metrics for evaluating them, and the datasets that underpin this research. Our synthesis reveals a growing demand for more rigorous, causal evaluation. By outlining the state of the art and identifying key challenges, this survey provides a roadmap for future research toward making models more transparent.

可解释性大模型概念描述神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。