对比了神经网络与图模型的语义空间结构,发现后者更清晰可读。
Geometry of Semantic Space: Comparative Study of Discrete and Continuous Models

- 用图结构建模语义关系,比传统向量嵌入更直观。
- 在法国全民辩论语料上,两者局部结构相似但整体拓扑差异大。
- 适合关注模型可解释性与语义结构的研究者。
本文研究自然语言模型背后的语义几何结构。比较了监督式向量嵌入(如CamemBERT)与基于词汇共现的图模型,后者更直接编码语义关系。尽管基于Transformer的嵌入表现优异,其诱导出的几何分布常不理想;而图模型则展现出更清晰、更符合人类认知的意义组织。我们提出一种方法,可基于图结构或嵌入拓扑进行对比分析。实验基于法语“全民大辩论”语料库(包含公民公共讨论文本),结果显示两者具有相似的局部拓扑,但整体结构和拓扑显著不同。这些发现表明,深度监督模型与图模型视角互补,为引导神经架构向更稳定、可解释的收敛方向发展提供了新路径。
原文摘要 · Abstract (English)
This work examines the semantic geometry underlying NLP models. We compare supervised vector embeddings, such as CamemBERT, with lexical co-occurrence graphs that encode semantic relations more directly. While transformer-based embeddings achieve strong performance, their induced geometries often display unsatisfactory distributions. In contrast, graph-based models reveal a clearer and more human-readable organization of meaning. We have implemented a methodology that allows us to perform a comparative analysis either based on the structure of the graphs or based on the topology of the embeddings induced by these two approaches. The results of the comparison -- applied to the French "Great National Debate" corpus a collection of citizen contributions to the public debate -- show a similar local topology but a very different overall structure and topology. Theses findings suggest complementary perspectives between deep supervised models and graph-based models, considering a new pathway to guide neural architectures toward more stable and interpretable convergence with graphs structures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。