arXiv:2508.13729cs.CLcs.AI2025-08被引 2

揭示词向量解释力的局限性:预测准确不等于真正理解语义。

Prediction is not Explanation: Revisiting the Explanatory Capacity of Mapping Embeddings

  • 通过映射向量空间与语义特征,检验词向量的可解释性
  • 发现即使随机信息也能被高精度预测,说明结果受算法上限主导
  • 提醒研究者勿仅凭预测性能判断语义捕捉能力,适合模型可解释性研究者

理解深度学习模型中隐含知识的编码方式,对提升人工智能系统的可解释性至关重要。本文考察了常用于解释词向量(词嵌入)中编码知识的方法,这些方法通常将嵌入映射到人类可读的语义特征集合(即特征规范)。以往工作假设:若能从词向量准确预测这些语义特征,则表明向量中包含相应知识。本文挑战这一假设,证明仅靠预测准确率无法可靠反映真实的基于特征的可解释性。我们发现,即使随机生成的信息也能被成功预测,说明结果主要由算法的理论上限决定,而非词向量中真正的语义表征。因此,仅依据预测性能比较不同数据集,无法可靠判断哪个数据集更被词向量所捕获。分析表明,这类映射更多反映向量空间中的几何相似性,而非语义属性的真实涌现。

原文摘要 · Abstract (English)

Understanding what knowledge is implicitly encoded in deep learning models is essential for improving the interpretability of AI systems. This paper examines common methods to explain the knowledge encoded in word embeddings, which are core elements of large language models (LLMs). These methods typically involve mapping embeddings onto collections of human-interpretable semantic features, known as feature norms. Prior work assumes that accurately predicting these semantic features from the word embeddings implies that the embeddings contain the corresponding knowledge. We challenge this assumption by demonstrating that prediction accuracy alone does not reliably indicate genuine feature-based interpretability. We show that these methods can successfully predict even random information, concluding that the results are predominantly determined by an algorithmic upper bound rather than meaningful semantic representation in the word embeddings. Consequently, comparisons between datasets based solely on prediction performance do not reliably indicate which dataset is better captured by the word embeddings. Our analysis illustrates that such mappings primarily reflect geometric similarity within vector spaces rather than indicating the genuine emergence of semantic properties.

可解释性词向量语义分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。