arXiv:2602.09992cs.CLcs.AI2026-02被引 1

神经语言模型仅靠少量数据就能学会语法,但效率远不如孩子。

A Unified Assessment of the Poverty of the Stimulus Argument for Neural Language Models

  • 构建统一基准测试,覆盖四种经典语法现象
  • 1000万词训练下模型仍能超越随机水平的语法泛化
  • 人类特有的学习效率依赖当前模型未实现的先验知识

近期研究利用人工神经网络(ANN)评估了刺激贫乏假说(PoSH)。结果表明,基于ANN的语言模型可在缺乏传统语言学家假设的结构归纳偏置情况下,从有限输入中习得某些结构依赖性泛化。然而,现有研究多聚焦单一现象且评价协议各异,难以判断结论是否具有普适性。本文提出 poshbench,一个涵盖四种经典PoS现象的统一基准。在Transformer、LSTM和n-gram模型上进行训练,发现神经网络模型在仅1000万词输入下即可实现高于随机水平的泛化能力,但随着输入规模增长,其学习效率仍显著低于儿童。此外,受认知启发的归纳偏置虽显著提升整体句法能力,却未一致促进与PoS相关的泛化。研究为‘先天语法是唯一泛化路径’的主张提供了系统性质疑,同时暗示人类级学习效率需超越当前模型所采用的归纳偏置。

原文摘要 · Abstract (English)

Several recent contributions have evaluated the Poverty of the Stimulus Hypothesis (PoSH) using Artificial Neural Networks (ANNs). The results suggest that ANN-based language models can acquire certain structure-dependent generalizations from limited input without structural inductive biases of the type traditionally hypothesized by linguists. However, existing studies have largely focused on individual phenomena and adopted different evaluation protocols, leaving it unclear whether previous findings generalize across phenomena and learning conditions. We introduce \poshbench, a unified benchmark covering four canonical PoS phenomena. Training Transformer, LSTM, and n-gram models, we find that ANN-based models can achieve above-chance generalization from surprisingly limited input (10M words), but they show less efficient learning than children as input scale grows. Moreover, cognitively motivated inductive biases substantially improve broad syntactic competence but do not consistently translate to PoS-relevant generalization. Our findings provide systematic evidence challenging the claim that innate syntax is the only possible route to generalization, while suggesting that human-like learning efficiency requires inductive biases beyond those implemented here.

语言模型语法习得神经网络认知科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。