arXiv:2503.03666cs.CLcs.LG2025-03被引 14

发现大模型内部存在可识别的概念向量,但抽象概念表征有限。

Analogical Reasoning Inside Large Language Models: Concept Vectors and the Limits of Abstraction

  • 通过注意力头定位不变的概念向量,作为语义特征探测器。
  • 概念向量能独立于输出正确形成,但模型仍可能答错。
  • 对具体概念有效,抽象概念如'前/后'则无稳定线性表征。

类比推理依赖概念抽象,但大语言模型是否具备此类内部表示尚不明确。我们分析了从模型激活中提取的压缩表示,发现用于上下文学习的任务功能向量(FVs)并不对输入形式变化(如开放问答与多选题)保持不变,表明其不仅编码概念。通过表示相似性分析(RSA),我们定位到少数注意力头编码出对'反义词'等词汇概念保持不变的概念向量(CVs)。这些CVs作为独立于最终输出的特征探测器工作——模型可能形成正确内部表征却仍输出错误答案。此外,CVs可被用来因果引导模型行为。但对于'前'、'后'等更抽象概念,未观察到稳定的线性表示,这与模型在这些领域表现出的泛化能力不足相关。

原文摘要 · Abstract (English)

Analogical reasoning relies on conceptual abstractions, but it is unclear whether Large Language Models (LLMs) harbor such internal representations. We explore distilled representations from LLM activations and find that function vectors (FVs; Todd et al., 2024) - compact representations for in-context learning (ICL) tasks - are not invariant to simple input changes (e.g., open-ended vs. multiple-choice), suggesting they capture more than pure concepts. Using representational similarity analysis (RSA), we localize a small set of attention heads that encode invariant concept vectors (CVs) for verbal concepts like "antonym". These CVs function as feature detectors that operate independently of the final output - meaning that a model may form a correct internal representation yet still produce an incorrect output. Furthermore, CVs can be used to causally guide model behaviour. However, for more abstract concepts like "previous" and "next", we do not observe invariant linear representations, a finding we link to generalizability issues LLMs display within these domains.

类比推理概念表征大模型机制注意力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。