arXiv:2509.19269cs.CL2025-09EMNLP

用原型描述提取大模型中的概念空间,让AI解释更直观。

Extracting Conceptual Spaces from LLMs Using Prototype Embeddings

  • 用原型描述生成嵌入向量,对应概念维度如‘甜度’。
  • 微调后原型嵌入与概念空间维度对齐,效果显著提升。
  • 适合想理解大模型内部表征的解释性研究者。

概念空间通过认知有意义的维度(如感知特征)表示实体和概念,广泛应用于认知科学,并有望成为可解释AI的基础。然而,学习概念空间极为困难,尽管近期的大语言模型(LLMs)已能捕捉到显著的感知特征。目前仍缺乏有效的提取方法。现有技术可提取LLM嵌入,但难以编码底层特征。本文提出一种策略:将特征(如甜度)通过其原型描述(如极甜食物)的嵌入来编码。为进一步优化,我们微调LLM,使原型嵌入与概念空间维度对齐。实证分析表明该方法高度有效。

原文摘要 · Abstract (English)

Conceptual spaces represent entities and concepts using cognitively meaningful dimensions, typically referring to perceptual features. Such representations are widely used in cognitive science and have the potential to serve as a cornerstone for explainable AI. Unfortunately, they have proven notoriously difficult to learn, although recent LLMs appear to capture the required perceptual features to a remarkable extent. Nonetheless, practical methods for extracting the corresponding conceptual spaces are currently still lacking. While various methods exist for extracting embeddings from LLMs, extracting conceptual spaces also requires us to encode the underlying features. In this paper, we propose a strategy in which features (e.g. sweetness) are encoded by embedding the description of a corresponding prototype (e.g. a very sweet food). To improve this strategy, we fine-tune the LLM to align the prototype embeddings with the corresponding conceptual space dimensions. Our empirical analysis finds this approach to be highly effective.

概念空间大模型解释原型嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。