arXiv:2410.00153cs.CLcs.AI2024-10ICLR被引 25

用高斯分布建模概念子空间,提升大模型语义理解的稳定性

Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution

  • 将单一概念向量扩展为高斯分布的子空间,增强鲁棒性
  • 在多模型上验证了子空间的忠实性与合理性,效果优于传统方法
  • 适用于情感控制等实际应用,兼顾生成流畅性与可控性

探测大语言模型(LLMs)中学习到的概念对理解其内部语义编码至关重要。传统的线性分类器方法虽能定位某一概念的向量表示,但该向量受数据和训练过程影响显著,导致结果不稳定,限制了其在真实场景中的应用。为此,我们提出一种新方法,通过线性探针扩展概念向量为高斯概念子空间(Gaussian Concept Subspace, GCS),以更准确地表征概念的内在结构。我们在多个不同规模与架构的LLM上验证了GCS在忠实性与合理性方面的有效性,并通过表示干预任务展示了其在情感控制等实际应用中的性能优势。实验表明,使用GCS进行概念操控时,在保持自然语言生成流畅性的前提下,能有效提升控制效果。

原文摘要 · Abstract (English)

Probing learned concepts in large language models (LLMs) is crucial for understanding how semantic knowledge is encoded internally. Training linear classifiers on probing tasks is a principle approach to denote the vector of a certain concept in the representation space. However, the single vector identified for a concept varies with both data and training, making it less robust and weakening its effectiveness in real-world applications. To address this challenge, we propose an approach to approximate the subspace representing a specific concept. Built on linear probing classifiers, we extend the concept vectors into Gaussian Concept Subspace (GCS). We demonstrate GCS's effectiveness through measuring its faithfulness and plausibility across multiple LLMs with different sizes and architectures. Additionally, we use representation intervention tasks to showcase its efficacy in real-world applications such as emotion steering. Experimental results indicate that GCS concept vectors have the potential to balance steering performance and maintaining the fluency in natural language generation tasks.

概念探测高斯分布语言模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。