测试大模型对语法形式与语义关系的理解能力,发现其仍存在明显短板。
CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
- 基于构式语法设计最小对比对,评估模型对形式-意义配对的理解
- 大型模型在九类句式中仍无法稳定识别语义关系,理解能力发展缓慢
- 适合研究语言模型语义整合能力及学习路径的学者使用
近期研究从语言学视角探讨语言模型如何习得语言,但多数基准测试仅关注语法正确性,对语法形式传达语义的能力关注不足。我们提出语言模型构式理解评估的最小对比对基准(CxMP),基于构式语法,将形式-意义配对视为基本语言单位。CxMP通过九种构式类型(包括let-alone、 caused motion、 ditransitive等)的受控最小对比对设计,评估模型对构式隐含语义关系的解读能力。结果表明,尽管句法能力早期即显现,但构式理解呈渐进式发展,即便在大型语言模型中仍存在显著局限。这揭示了语言模型在形式与意义整合上的持续差距,为研究模型的构式理解与学习轨迹提供了框架。
原文摘要 · Abstract (English)
Recent work has examined language models from a linguistic perspective to better understand how they acquire language. Most existing benchmarks focus on judging grammatical acceptability, whereas the ability to interpret meanings conveyed by grammatical forms has received much less attention. We introduce the Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models (CxMP), a benchmark grounded in Construction Grammar that treats form-meaning pairings, or constructions, as fundamental linguistic units. CxMP evaluates whether models can interpret the semantic relations implied by constructions, using a controlled minimal-pair design across nine construction types, including the let-alone, caused motion, and ditransitive constructions. Our results show that while syntactic competence emerges early, constructional understanding develops more gradually and remains limited even in large language models (LLMs). CxMP thus reveals persistent gaps in how language models integrate form and meaning, providing a framework for studying constructional understanding and learning trajectories in language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。