arXiv:2603.15295cs.CLcs.DB2026-03中稿 · LREC 2026被引 1

构建跨语言动词交替数据集,测试大模型对句间语法模式的理解能力

Datasets for Verb Alternations across Languages: BLM Templates and Data Augmentation Strategies

  • 设计黑鸟语言矩阵(BLM)任务,通过语法规则匹配句式模式
  • 涵盖英德意希四种语言,包含数千个动词交替问题,验证模型系统性知识
  • 结合合成与自然数据增强策略,适合作为语言模型评估基准

大语言模型在各类句子级语言现象中表现优异,但对其跨句子的系统性句法模式(如英语、德语、意大利语中的状态变化和宾语省略结构,希伯来语的词干结构)的理解仍缺乏研究。本文构建了针对四种语言的归纳式数据集,包含数千个黑鸟语言矩阵(BLM)问题。BLM任务是一种专为语言设计的类RPM/ARC谜题,要求模型根据句法和语义规则选择完成特定模式的句子。我们提出三种不同复杂度的模板,并在合成与自然数据上应用语言学驱动的数据增强策略。在英语、意大利语、德语和希伯来语上提供简单基线结果,证明该数据集具有诊断价值。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown remarkable performance across various sentence-based linguistic phenomena, yet their ability to capture cross-sentence paradigmatic patterns, such as verb alternations, remains underexplored. In this work, we present curated paradigm-based datasets for four languages, designed to probe systematic cross-sentence knowledge of verb alternations (change-of-state and object-drop constructions in English, German and Italian, and Hebrew binyanim). The datasets comprise thousands of the Blackbird Language Matrices (BLMs) problems. The BLM task -- an RPM/ARC-like task devised specifically for language -- is a controlled linguistic puzzle where models must select the sentence that completes a pattern according to syntactic and semantic rules. We introduce three types of templates varying in complexity and apply linguistically-informed data augmentation strategies across synthetic and natural data. We provide simple baseline performance results across English, Italian, German, and Hebrew, that demonstrate the diagnostic usefulness of the datasets.

语言模型动词交替数据集多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。