arXiv:2501.04661cs.CLcs.AI2025-01被引 6

测试大模型对少见语义构造的理解能力,发现其泛化能力远不如人类。

Beyond Memorization: Assessing Semantic Generalization in Large Language Models Using Phrasal Constructions

  • 用构式语法设计诊断性评测,检验模型对非常见句式的理解。
  • 顶尖模型在相似句式不同语义任务上性能下降超40%。
  • 适合研究模型语义泛化与人类语言认知差异的学者参考。

大规模预训练数据带来了评估挑战:如何区分模型对常见语言现象的机械记忆与对真实世界中少见语言模式的泛化能力。为此,我们基于构式语法(CxG)构建了一种诊断性评估方法。CxG提供心理学基础框架,将句法形式与抽象语义直接关联。我们的新评测数据集包含英语短语构造,这些构造虽在预训练数据中不常见,但人类能轻松抽象并生成创造性实例。该数据集旨在回答两个核心问题:一是模型能否理解预训练数据中罕见但人类易懂的语义;二是当句法相同但语义不同,模型是否能正确使用相应构式语义。结果表明,包括GPT-o1在内的先进模型在第二项任务中性能下降超过40%,暴露出其无法像人类一样根据句法一致形式推导出不同构式意义。我们已公开数据集及实验数据(含提示和模型响应)。

原文摘要 · Abstract (English)

The web-scale of pretraining data has created an important evaluation challenge: to disentangle linguistic competence on cases well-represented in pretraining data from generalization to out-of-domain language, specifically the dynamic, real-world instances less common in pretraining data. To this end, we construct a diagnostic evaluation to systematically assess natural language understanding in LLMs by leveraging Construction Grammar (CxG). CxG provides a psycholinguistically grounded framework for testing generalization, as it explicitly links syntactic forms to abstract, non-lexical meanings. Our novel inference evaluation dataset consists of English phrasal constructions, for which speakers are known to be able to abstract over commonplace instantiations in order to understand and produce creative instantiations. Our evaluation dataset uses CxG to evaluate two central questions: first, if models can 'understand' the semantics of sentences for instances that are likely to appear in pretraining data less often, but are intuitive and easy for people to understand. Second, if LLMs can deploy the appropriate constructional semantics given constructions that are syntactically identical but with divergent meanings. Our results demonstrate that state-of-the-art models, including GPT-o1, exhibit a performance drop of over 40% on our second task, revealing a failure to generalize over syntactically identical forms to arrive at distinct constructional meanings in the way humans do. We make our novel dataset and associated experimental data, including prompts and model responses, publicly available.

大模型评估语义泛化构式语法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。