arXiv:2603.25227cs.CL2026-03中稿 · the Workshop on St…

用法语意大利语被动句对比,发现自然数据更利于模型掌握语言规律。

Comparing Natural and Synthetic Structured Data: A Study of the Passive Verb Alternation in French and Italian

  • 用真实和合成句子构建结构化数据集,测试模型对被动句模式的掌握。
  • 合成数据让模型表现达顶峰,但无法迁移至真实句子;自然数据训练的模型泛化更强。
  • 适合研究语言模型语法理解能力、评估方法设计的学者参考。

本研究比较了自然数据与合成数据对大语言模型(LLMs)训练与评估的影响,以法语和意大利语的被动动词交替现象为例。我们采用黑鸟语言矩阵(Blackbird Language Matrices, BLMs),一种用于探测句子集合中潜在语言模式的结构化数据集。将从通用依赖库(Universal Dependencies)提取的真实句子与合成句子分别嵌入结构化模板进行对比。实验表明,当模型在合成数据上训练并测试时可达到上限性能,但难以泛化到真实句子;而以自然数据训练的模型在真实与合成测试集上均表现稳健,显示出更强的抽象语言模式捕捉能力。结果证实了自然数据与结构化设置在语言模型句法与语义知识探测中的价值。

原文摘要 · Abstract (English)

This study compares the impact of natural and synthetic data on training and evaluating large language models (LLMs), using the case of passive verb alternation in French and Italian. We use Blackbird Language Matrices (BLMs), structured datasets designed to probe linguistic knowledge of underlying patterns across sentence sets. We compare structured templates instantiated with natural sentences extracted from Universal Dependencies to structured templates of synthetic sentences. Experiments show that while models achieve ceiling performance when trained and tested on synthetic datasets, they do not reliably generalize to natural sentences. In contrast, models trained on natural data exhibit robust performance across both natural and synthetic test suites, demonstrating their superior ability to capture abstract linguistic patterns. These results corroborate the value of natural data and of structured set ups in linguistic evaluation for probing LLMs' syntactic and semantic knowledge.

语言模型语法分析数据对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。