研究语言模型语义空间中关系几何的表示能力,发现不同关系的几何特征差异显著。
Relation Geometry in Semantic Space of Language Models

- 从三方面分析语义空间中词语关系的几何分布
- 不对称关系在空间中占据更清晰区域,属性编码优于对称关系
- 上下文信息对掩码与扩散模型影响更大,词汇信息对因果模型更重要
当前语言模型在生成词向量表示方面已取得高质量成果,但其语义空间中语义关系的几何表征程度尚不明确。本文从三个角度研究此类语义空间的关系几何:首先考察特定目标词的相关词(称为相关项)是否占据语义空间中的同一区域,以及不同关系对应的区域是否可区分;其次验证语义空间是否反映关系的对称性、非对称性和传递性等特性;最后探讨目标词与相关项的表面形式或上下文信息,哪种对关系几何影响更大。实验使用六种语义关系,在因果、掩码和扩散语言模型上进行。结果表明,非对称关系的相关项在语义空间中相对清晰地占据独立区域,其属性编码程度中等但优于对称关系。在所评估模型中,词汇信息对因果模型影响更强,而上下文信息对掩码和扩散模型更具影响力。实证显示,并非所有语义关系在语义空间中都同等良好地被表示,提示仅靠分布信息学习某些关系可能存在局限。
原文摘要 · Abstract (English)
When it comes to generating vector representations of words, current language models are achieving high-quality results. However, what is not known is the extent to which knowledge about semantic relations is represented in the geometry of the semantic spaces created in this way. In order to answer this question, we study the relation geometry of such semantic spaces from three perspectives. We first examine whether words standing in a particular relation to a target word~(called relata) occupy the same region in semantic space, and whether the regions corresponding to different relations are distinct from each other. We then verify to what extent semantic spaces reflect certain well-known properties of relations, such as symmetry, asymmetry, and transitivity. Finally, we consider which information about the target words and relata is more important for relation geometry: their surface forms, or their contexts. We conduct experiments on six semantic relations using causal, masked, and diffusion language models. The results show that relata in asymmetric relations relatively clearly occupy a distinct region in semantic space. Asymmetric relations' properties are only moderately well encoded in the semantic space, yet better than those of symmetric ones. Furthermore, when considering the question which information source has the strongest impact on results amongst the models we evaluated, we find that lexical information tends to be more important for the causal language model, whereas contextual information is more important for the masked and diffusion language models. Our results empirically show that relation geometry is not equally well-represented for all relations in semantic space, suggesting that there is a difference in how well semantic relations might be learned from distributional information alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。