提出新方法从大模型中提取可解释概念,理论更扎实。
Concept Component Analysis: A Principled Approach for Concept Extraction in LLMs
- 基于隐变量模型,将模型激活看作概念后验的线性组合。
- 在多个大模型上成功提取出有意义的概念,优于传统方法。
- 适合关注模型可解释性的研究者和应用开发者。
开发人类可理解的大语言模型(LLMs)解释方法对关键领域部署至关重要。机制可解释性通过从模型激活中提取人类可理解的过程与概念来缓解这一问题。稀疏自编码器(SAEs)已成为一种流行方法,通过将模型内部表示分解为字典来提取可解释且单义的概念。尽管取得了经验进展,但SAEs存在根本性理论模糊:模型表示与人类可理解概念之间的明确对应关系尚不清晰。这种缺乏理论基础导致方法设计和评估标准上的诸多挑战。本文表明,在温和假设下,可通过潜在变量模型视角,将LLM表示近似为给定输入上下文的概念后验的线性组合。这启发了一种原则性概念提取框架——概念组件分析(ConCA),旨在通过无监督线性解混过程从LLM表示中恢复每个概念的对数后验。我们探索了具体变体——稀疏ConCA,利用稀疏性先验解决解混问题的固有不适定性。我们实现了12种稀疏ConCA变体,并在多个大模型上展示了其提取有意义概念的能力,相较于SAEs具有理论支持的优势。
原文摘要 · Abstract (English)
Developing human understandable interpretation of large language models (LLMs) becomes increasingly critical for their deployment in essential domains. Mechanistic interpretability seeks to mitigate the issues through extracts human-interpretable process and concepts from LLMs' activations. Sparse autoencoders (SAEs) have emerged as a popular approach for extracting interpretable and monosemantic concepts by decomposing the LLM internal representations into a dictionary. Despite their empirical progress, SAEs suffer from a fundamental theoretical ambiguity: the well-defined correspondence between LLM representations and human-interpretable concepts remains unclear. This lack of theoretical grounding gives rise to several methodological challenges, including difficulties in principled method design and evaluation criteria. In this work, we show that, under mild assumptions, LLM representations can be approximated as a {linear mixture} of the log-posteriors over concepts given the input context, through the lens of a latent variable model where concepts are treated as latent variables. This motivates a principled framework for concept extraction, namely Concept Component Analysis (ConCA), which aims to recover the log-posterior of each concept from LLM representations through a {unsupervised} linear unmixing process. We explore a specific variant, termed sparse ConCA, which leverages a sparsity prior to address the inherent ill-posedness of the unmixing problem. We implement 12 sparse ConCA variants and demonstrate their ability to extract meaningful concepts across multiple LLMs, offering theory-backed advantages over SAEs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。