解析BERT向量中编码语言特征的特定维度,揭示哪些特征由固定维度承载。
Disentangling Linguistic Features with Dimension-Wise Analysis of Vector Embeddings
- 构建10类语言特征句对数据集,分离出同义、否定、时态等关键特征。
- 发现否定和极性在特定维度稳定编码,同义等特征模式更复杂。
- 提出EDI评分量化维度对语言属性的影响,助力模型可解释性提升。
理解神经嵌入(如BERT)的内部机制仍具挑战,因其高维且不透明。本文提出框架,识别向量嵌入中编码具体语言属性(LPs)的特定维度。我们构建了包含10个关键语言特征(如同义、否定、时态、数量)的语义区分句对数据集LDSP-10。通过威尔科克森符号秩检验、互信息和递归特征消除等方法分析BERT嵌入,识别每类语言属性最显著的维度。引入新指标嵌入维度影响(EDI)分数,量化各维度对语言属性的相关性。研究发现,否定和极性等属性在特定维度上稳健编码,而同义性等则呈现更复杂的分布模式。该工作深化了对嵌入可解释性的理解,为开发更透明、高效的语言模型提供依据,有助于缓解模型偏见并推动人工智能负责任部署。
原文摘要 · Abstract (English)
Understanding the inner workings of neural embeddings, particularly in models such as BERT, remains a challenge because of their high-dimensional and opaque nature. This paper proposes a framework for uncovering the specific dimensions of vector embeddings that encode distinct linguistic properties (LPs). We introduce the Linguistically Distinct Sentence Pairs (LDSP-10) dataset, which isolates ten key linguistic features such as synonymy, negation, tense, and quantity. Using this dataset, we analyze BERT embeddings with various methods, including the Wilcoxon signed-rank test, mutual information, and recursive feature elimination, to identify the most influential dimensions for each LP. We introduce a new metric, the Embedding Dimension Impact (EDI) score, which quantifies the relevance of each embedding dimension to a LP. Our findings show that certain properties, such as negation and polarity, are robustly encoded in specific dimensions, while others, like synonymy, exhibit more complex patterns. This study provides insights into the interpretability of embeddings, which can guide the development of more transparent and optimized language models, with implications for model bias mitigation and the responsible deployment of AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。