arXiv:2506.06169cs.CLcs.AI2025-06被引 3

用可解释语义空间分析语言模型对双宾与介词结构的语义差异。

semantic-features: A User-Friendly Tool for Studying Contextual Word Embeddings in Interpretable Semantic Spaces

  • 将上下文词向量投影到可解释语义空间,分析句法结构影响
  • 3个掩码语言模型均显示预期语义敏感性差异
  • 适合语言学研究者探索模型语义理解机制

我们介绍 semantic-features,一个基于 Chronis 等人(2023)的可扩展、易用的库,用于通过将语言模型的上下文词向量投影到可解释空间来研究其语义表示。我们在一项实验中考察了间接宾语结构(介词型或双宾型)的选择如何影响语义解读(Bresnan, 2007)。具体而言,测试“我寄了信给伦敦”中“伦敦”被理解为有生命实体(如人名)的可能性是否高于“我寄了伦敦这封信”。为此,我们构建了一个包含450组句子对的数据集,每组包含两种间接宾语结构,且接收者在人与地点身份上具有歧义。通过应用 semantic-features,我们发现三个掩码语言模型的上下文词向量均表现出预期的语义敏感性。这使我们对工具的实用性充满信心。

原文摘要 · Abstract (English)

We introduce semantic-features, an extensible, easy-to-use library based on Chronis et al. (2023) for studying contextualized word embeddings of LMs by projecting them into interpretable spaces. We apply this tool in an experiment where we measure the contextual effect of the choice of dative construction (prepositional or double object) on the semantic interpretation of utterances (Bresnan, 2007). Specifically, we test whether "London" in "I sent London the letter." is more likely to be interpreted as an animate referent (e.g., as the name of a person) than in "I sent the letter to London." To this end, we devise a dataset of 450 sentence pairs, one in each dative construction, with recipients being ambiguous with respect to person-hood vs. place-hood. By applying semantic-features, we show that the contextualized word embeddings of three masked language models show the expected sensitivities. This leaves us optimistic about the usefulness of our tool.

词向量语义分析语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。