arXiv:2412.15698cs.LG2024-12

提出概念边界向量,更精准捕捉模型潜空间中的语义概念。

Concept Boundary Vectors

  • 从概念的潜空间边界构造向量,揭示语义差异方向。
  • 实验证明其能更准确表征概念语义,优于传统激活向量。
  • 适合想理解模型内部表示的研究者与可解释性应用者。

机器学习模型通常以简单目标(如下一个词预测)进行训练,但在部署时却表现出对输入数据更本质的表征能力。理解这些表征有助于解释模型输出并提升其语义显著性。概念向量旨在将输入数据中的概念映射到模型潜空间中的方向向量。本文提出概念边界向量,一种基于概念潜空间表示边界的向量构造方法。实证表明,该方法能有效捕捉概念的语义含义,并在效果上优于传统的概念激活向量(Concept Activation Vectors, CAV)。

原文摘要 · Abstract (English)

Machine learning models are trained with relatively simple objectives, such as next token prediction. However, on deployment, they appear to capture a more fundamental representation of their input data. It is of interest to understand the nature of these representations to help interpret the model's outputs and to identify ways to improve the salience of these representations. Concept vectors are constructions aimed at attributing concepts in the input data to directions, represented by vectors, in the model's latent space. In this work, we introduce concept boundary vectors as a concept vector construction derived from the boundary between the latent representations of concepts. Empirically we demonstrate that concept boundary vectors capture a concept's semantic meaning, and we compare their effectiveness against concept activation vectors.

概念向量潜空间可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。