arXiv:2605.23710cs.CL2026-05

用图模型分析词嵌入中语义类型匹配与强制现象

A graph-based analysis of semantic types and coercion in contextualized word embeddings

  • 构建词嵌入的语义类型图,通过邻居类型分布分析
  • 感知增强嵌入能更好反映语义类型信息,区分匹配与强制句
  • 适合研究语言理解、词向量语义表征的学者参考

名词与其上下文之间的语义类型不匹配是强制现象的核心。本文提出一种基于图的方法,分析词汇和上下文类型信息在词嵌入中的体现。选取十个语义类型的名词,标注语料库实例的类型匹配情况(匹配、强制、其他不匹配、无限制),并使用BERT和感知增强嵌入构建图结构。提出两个指标——邻居类型概率(NTP)和邻居类型熵(NTE)——用于分析邻域类型分布。结果表明,使用感知增强嵌入构建的图能更准确反映语义类型信息,且可通过所提指标区分匹配句与不匹配句。

原文摘要 · Abstract (English)

Semantic type mismatch between a noun and its context is central to coercion phenomena. This paper introduces a graph-based method to examine how lexical and contextual type information is reflected in word embeddings. We select nouns from ten semantic types, annotate corpus instances for type matching (matching vs. coercion vs. other mismatch vs. unrestricted), and construct graphs using BERT and sense-enhanced embeddings. Two metrics -- Neighbor Type Probability (NTP) and Neighbor Type Entropy (NTE) -- are proposed to analyze neighborhood type distributions. Results show that graphs constructed with sense-enhanced embeddings reflect semantic type information better, and matching and mismatch sentences can be distinguished through the proposed metrics.

词嵌入语义类型图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。