arXiv:2412.01131cs.CL2024-12被引 1

对比大模型与人类在五类语义关系上的理解差异,发现模型普遍存在知识短板。

A Comprehensive Evaluation of Semantic Relation Knowledge of Pretrained Language Models and Humans

  • 构建涵盖五类语义关系的评估框架,引入五项新指标。
  • 六种模型在所有关系上均显著落后于人类,仅反义关系表现较好。
  • 适合研究模型语义理解能力或人机认知差异的学者参考。

近期大量研究关注预训练语言模型(PLMs)对语言不同层面的理解内容及其学习机制。其中一类研究聚焦于模型对语义关系的知识掌握情况,但以往工作仅关注单一关系——上下位关系(hypernymy),且未在相同任务下衡量人类表现。这导致当前对模型语义关系知识的了解仍不完整。为此,本文提出一个综合性评估框架,涵盖五类超越上下位关系的语义关系:下位关系(hyponymy)、整体-部分关系(holonymy)、组成部分关系(meronymy)、反义关系(antonymy)和同义关系(synonymy)。我们采用五项度量标准(其中两项为新提出)来评估模型在一致性、完备性、对称性、典型性和可区分性方面的表现。通过六个PLMs(四个掩码模型、两个因果模型)的广泛实验,结果表明:所有语义关系上,模型的表现均显著低于人类;尽管因果模型应用广泛,但其性能并不总优于掩码模型;反义关系是唯一例外,所有模型在此任务上表现尚可。评估数据集已开源:https://github.com/hancules/ProbeResponses。

原文摘要 · Abstract (English)

Recently, much work has concerned itself with the enigma of what exactly pretrained language models~(PLMs) learn about different aspects of language, and how they learn it. One stream of this type of research investigates the knowledge that PLMs have about semantic relations. However, many aspects of semantic relations were left unexplored. Generally, only one relation has been considered, namely hypernymy. Furthermore, previous work did not measure humans' performance on the same task as that performed by the PLMs. This means that at this point in time, there is only an incomplete view of the extent of these models' semantic relation knowledge. To address this gap, we introduce a comprehensive evaluation framework covering five relations beyond hypernymy, namely hyponymy, holonymy, meronymy, antonymy, and synonymy. We use five metrics (two newly introduced here) for recently untreated aspects of semantic relation knowledge, namely soundness, completeness, symmetry, prototypicality, and distinguishability. Using these, we can fairly compare humans and models on the same task. Our extensive experiments involve six PLMs, four masked and two causal language models. The results reveal a significant knowledge gap between humans and models for all semantic relations. In general, causal language models, despite their wide use, do not always perform significantly better than masked language models. Antonymy is the outlier relation where all models perform reasonably well. The evaluation materials can be found at https://github.com/hancules/ProbeResponses.

语义理解大模型评测认知对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。