arXiv:2602.02726cs.LGcs.CL2026-02

用向量量化方法高效学习语言模型隐状态中的语义概念。

Vector Quantized Latent Concepts: A Scalable Alternative to Clustering-Based Concept Discovery

  • 用向量量化在冻结的隐藏状态上学习离散语义概念
  • 计算成本接近K-Means,比层次聚类更可扩展,且语义连贯性更好
  • 适合需要高效解释大模型内部表征的研究者

大型语言模型(LLMs)在其隐藏状态中编码了丰富的语义信息,但理解这些内部表示捕捉的内容仍具挑战。从隐藏状态中提取的潜在概念为解释LLMs提供了有前景的方向,但现有基于聚类的方法存在权衡:层次聚类生成语义一致的概念,但因二次内存开销仅限于小数据集;K-Means计算高效,但可能产生语义连贯性较弱的概念。我们提出向量量化潜在概念(VQLC),一种在冻结隐藏状态上学习潜在概念代码本的离散概念学习框架。在12个数据集-模型组合中,VQLC保持与K-Means相近的计算成本,比层次聚类更具可扩展性,并在忠实度上保持竞争力,尤其在解码器单体模型上表现更优。基于LLM的评估、定性分析及与稀疏自编码器(SAE)的对比表明,所学概念具有可解释性且任务相关。

原文摘要 · Abstract (English)

Large language models (LLMs) encode rich semantic information in their hidden states, yet it remains difficult to understand what information these internal representations capture. Latent concepts extracted from hidden states offer a promising direction for interpreting LLMs, but existing clustering-based methods face a trade-off: hierarchical clustering produces coherent concepts but is limited to small datasets due to its quadratic memory cost, while K-Means scales efficiently but may yield less semantically coherent concepts. We propose Vector Quantized Latent Concept (VQLC), a discrete concept learning framework that learns a codebook of latent concepts on frozen hidden states. Across 12 dataset-model settings, VQLC stays close to K-Means in computational cost, scales better than hierarchical clustering, and remains competitive in faithfulness, with the clearest gains on decoder-only models. LLMs-based evaluation, qualitative analysis, and a Sparse Autoencoder (SAE) comparison demonstrate that the learned concepts are interpretable and task-relevant.

概念发现向量量化模型解释大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。