arXiv:2605.28616cs.CLcs.AI2026-05

用儿童语言习得标准评估大模型的句法与语用能力

Measuring Form and Function in Language Models

论文配图:Measuring Form and Function in Language Models
图 1 · 摘自论文原文
  • 设计新提示方法CAC,精准测试模型句法与语用知识
  • 现有模型均未同时达标人类儿童的语法和语用表现
  • 适合关注模型认知能力与语言理解深度的研究者

我们引入量化指标,借鉴儿童语言习得研究方法评估语言模型。聚焦英语中限定词的正式句法与功能语用特性,这些特征是幼儿早期准确掌握的内容。提出上下文替代选择(CAC)新提示方法,可针对性测试模型的句法与语用知识,并实现模型与儿童、以及独立建立的统计基准之间的直接比较。目前在相似数据量下训练的任何模型均未能同时达到人类儿童的正式与功能双重基准,但部分超大规模模型接近该水平。研究贡献在于方法与技术层面,重点强调语言模型的认知状态。

原文摘要 · Abstract (English)

We introduce quantitative metrics for child language acquisition to evaluate language models. Our focus is on the formal syntactic and functional discourse properties of determiners in English, which young children acquire early and accurately. We propose Contextual Alternative Choice (CAC), a new prompting method which provides targeted tests for both syntactic and discourse knowledge of language. The method enables direct comparison of language models against children, and more importantly, against statistical benchmarks independently established in empirical research. No current model trained on a comparable amount of data simultaneously meet both formal and functional benchmarks like human children, but some very large models do. We present our results as methodological and technical contributions, with specific emphasis on cognitive status of language models.

语言模型认知评估句法分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。