arXiv:2509.16422cs.CL2025-09被引 3

用构造语法评测大模型的深层语义理解能力,发现抽象模式下性能下降明显。

Evaluating CxG Generalisation in LLMs via Construction-Based NLI Fine Tuning

  • 基于模板生成合成数据,结合模型筛选构建8万句对抗性推理任务
  • 零样本测试显示大模型在抽象构造上准确率下降24%,最低达64%
  • 提供可扩展的评估框架,适合研究语言抽象与模型泛化者参考

我们探测大语言模型学习构造语法定义的深层形式-语义映射的能力。引入包含8万个句子的ConTest-NLI基准,覆盖从高度具体到高度抽象的八种英语构造。通过模板化生成与模型内过滤器结合的流水线,生成多样化的合成自然语言推理三元组,实现类人验证以保障挑战性和标签可靠性。对主流LLM进行零样本测试发现,自然数据(88%)与对抗数据(64%)间准确率下降24%,抽象构造最难处理。在部分ConTest-NLI上微调可提升最多9%性能,但结果仍揭示当前大模型存在持续的抽象能力鸿沟,并提供一种可扩展的构造感知学习评估框架。

原文摘要 · Abstract (English)

We probe large language models' ability to learn deep form-meaning mappings as defined by construction grammars. We introduce the ConTest-NLI benchmark of 80k sentences covering eight English constructions from highly lexicalized to highly schematic. Our pipeline generates diverse synthetic NLI triples via templating and the application of a model-in-the-loop filter. This provides aspects of human validation to ensure challenge and label reliability. Zero-shot tests on leading LLMs reveal a 24% drop in accuracy between naturalistic (88%) and adversarial data (64%), with schematic patterns proving hardest. Fine-tuning on a subset of ConTest-NLI yields up to 9% improvement, yet our results highlight persistent abstraction gaps in current LLMs and offer a scalable framework for evaluating construction-informed learning.

语言理解构造语法模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。