arXiv:2506.04373cs.CLcs.AI2025-06

用词级字典学习分解句向量,让语义结构可解释

Mechanistic Decomposition of Sentence Representations

  • 通过词级表示的字典学习,拆解句向量成分
  • 发现语义语法特征在向量中呈线性编码
  • 适合关注模型可解释性的研究者

句向量是现代自然语言处理和人工智能系统的核心,但其内部结构尚不明确。尽管可用余弦相似度等方法比较这些向量,但其贡献特征难以人工理解,嵌入内容因复杂的神经变换和最终的池化操作而变得不可追踪。为缓解此问题,我们提出一种新的机制分解方法,通过在词级表示上使用字典学习,将句向量分解为可解释的组件。我们分析了池化如何将这些特征压缩为句向量,并评估了句向量中潜藏的特征。该方法连接了词级机制可解释性与句级分析,使表示更透明可控。研究揭示了句向量空间中的多个有趣洞察,例如许多语义和句法特征在线性编码中存在。

原文摘要 · Abstract (English)

Sentence embeddings are central to modern NLP and AI systems, yet little is known about their internal structure. While we can compare these embeddings using measures such as cosine similarity, the contributing features are not human-interpretable, and the content of an embedding seems untraceable, as it is masked by complex neural transformations and a final pooling operation that combines individual token embeddings. To alleviate this issue, we propose a new method to mechanistically decompose sentence embeddings into interpretable components, by using dictionary learning on token-level representations. We analyze how pooling compresses these features into sentence representations, and assess the latent features that reside in a sentence embedding. This bridges token-level mechanistic interpretability with sentence-level analysis, making for more transparent and controllable representations. In our studies, we obtain several interesting insights into the inner workings of sentence embedding spaces, for instance, that many semantic and syntactic aspects are linearly encoded in the embeddings.

可解释性句向量字典学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。