arXiv:2411.02610cs.CLcs.AI2024-11被引 4

测试词向量模型对习语意义的理解能力,发现现有模型仍不准确。

Investigating Idiomaticity in Word Representations

  • 构建双语习语数据集,用最小差异句对比习语与字面意义
  • 多数模型看似相似实则无法捕捉习语真实含义,敏感度不足
  • 适合研究语言理解、语义表示或模型可解释性的研究人员

习语是人类语言的重要组成部分,常以压缩或约定的方式表达复杂概念(如“eager beaver”指积极热情的人)。然而其含义往往无法从单个成分的字面意义直接推导,这对组合性建模方法构成挑战。本文研究词向量模型在多大程度上能超越成分组合,捕捉名词复合词的习语性及其相关语义特性。聚焦英语和葡萄牙语中不同习语化程度的名词复合词,构建包含类型与实例级人工习语判断的最小差异句数据集,涵盖同义句与自然语境下的无义中性句子,共32,200句。提出细粒度的亲和度(Affinity)与缩放相似度(Scaled Similarity)指标,评估模型对影响习语性的微小变化的敏感性。多种代表性模型结果表明,尽管表面相似度高,当前模型仍未能准确表征习语性。不同上下文化水平模型的表现也显示,其上下文理解能力尚未能超越词汇表层线索,无法真正融入习语所需的语义信息。

原文摘要 · Abstract (English)

Idiomatic expressions are an integral part of human languages, often used to express complex ideas in compressed or conventional ways (e.g. eager beaver as a keen and enthusiastic person). However, their interpretations may not be straightforwardly linked to the meanings of their individual components in isolation and this may have an impact for compositional approaches. In this paper, we investigate to what extent word representation models are able to go beyond compositional word combinations and capture multiword expression idiomaticity and some of the expected properties related to idiomatic meanings. We focus on noun compounds of varying levels of idiomaticity in two languages (English and Portuguese), presenting a dataset of minimal pairs containing human idiomaticity judgments for each noun compound at both type and token levels, their paraphrases and their occurrences in naturalistic and sense-neutral contexts, totalling 32,200 sentences. We propose this set of minimal pairs for evaluating how well a model captures idiomatic meanings, and define a set of fine-grained metrics of Affinity and Scaled Similarity, to determine how sensitive the models are to perturbations that may lead to changes in idiomaticity. The results obtained with a variety of representative and widely used models indicate that, despite superficial indications to the contrary in the form of high similarities, idiomaticity is not yet accurately represented in current models. Moreover, the performance of models with different levels of contextualisation suggests that their ability to capture context is not yet able to go beyond more superficial lexical clues provided by the words and to actually incorporate the relevant semantic clues needed for idiomaticity.

语义表示习语理解词向量多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。