发现大模型中概念表征与任务驱动机制不同,概念向量更稳定。
Causality $\neq$ Invariance: Function and Concept Vectors in LLMs
- 用表示相似性分析找出跨格式稳定的概念向量
- 概念向量在多语言多题型下泛化能力更强
- 任务向量擅长特定格式,概念向量体现抽象表征
大型语言模型是否以抽象方式表征概念,即不依赖输入格式?我们重新考察了函数向量(FVs),这是一种紧凑的任务表征,能因果驱动上下文学习性能。在多个模型中,我们发现FVs并非完全不变:从不同输入格式(如开放问答与选择题)提取的FVs几乎正交,即使目标概念相同。我们识别出概念向量(CVs),其携带更稳定的概念表示。与FVs类似,CVs由注意力头输出组成;但不同于FVs,其头部通过表示相似性分析(RSA)筛选,仅保留跨格式一致编码概念的头。这些头虽出现在与FV相关头相似的层,但两者基本不重叠,暗示不同机制。操控实验表明,当提取与应用格式一致(如均为英文开放问答)时,FVs表现更优;而CVs在跨题型(开放问答与选择题)和跨语言场景中更具泛化能力。结果表明,大模型包含抽象概念表征,但其不同于驱动上下文学习的机制。
原文摘要 · Abstract (English)
Do large language models (LLMs) represent concepts abstractly, i.e., independent of input format? We revisit Function Vectors (FVs), compact representations of in-context learning (ICL) tasks that causally drive task performance. Across multiple LLMs, we show that FVs are not fully invariant: FVs are nearly orthogonal when extracted from different input formats (e.g., open-ended vs. multiple-choice), even if both target the same concept. We identify Concept Vectors (CVs), which carry more stable concept representations. Like FVs, CVs are composed of attention head outputs; however, unlike FVs, the constituent heads are selected using Representational Similarity Analysis (RSA) based on whether they encode concepts consistently across input formats. While these heads emerge in similar layers to FV-related heads, the two sets are largely distinct, suggesting different underlying mechanisms. Steering experiments reveal that FVs excel in-distribution, when extraction and application formats match (e.g., both open-ended in English), while CVs generalize better out-of-distribution across both question types (open-ended vs. multiple-choice) and languages. Our results show that LLMs do contain abstract concept representations, but these differ from those that drive ICL performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。