arXiv:2410.13259cs.CL2024-10被引 3

用人类语言习得路径评估大模型,发现其发展不完全相似。

From Babbling to Fluency: Evaluating the Evolution of Language Models in Terms of Human Language Acquisition

  • 构建三阶段框架,从词汇到复杂语法与逻辑推理评估模型
  • 新模型在词长、从句等易提取特征上表现接近人类,但进展有限
  • 训练数据特性影响模型能力,提示需关注数据多样性

我们从人类语言习得的视角审视语言模型的语言能力。基于经典语言发展理论,提出一个涵盖初步词汇理解、复杂语法和逻辑推理的三阶段评估框架,并采用语言学研究方法评估生成能力。结果表明,尽管近期语言模型整体性能优于早期模型,其发展轨迹并未严格遵循人类语言习得路径。在生成任务中,模型在信息易从语料中提取的方面(如平均词长、从句、助动词)更接近人类表现。然而,在从句和助动词等语料变异较小的维度上,新模型未表现出显著进步。注册理论可解释此现象:训练数据的语言特征对模型能力有显著影响。

原文摘要 · Abstract (English)

We examine the language capabilities of language models (LMs) from the critical perspective of human language acquisition. Building on classical language development theories, we propose a three-stage framework to assess the abilities of LMs, ranging from preliminary word understanding to complex grammar and complex logical reasoning. Using this framework, we evaluate the generative capacities of LMs using methods from linguistic research. Results indicate that although recent LMs outperform earlier models in overall performance, their developmental trajectory does not strictly follow the path of human language acquisition. Notably, in generation tasks, LMs are more similar to human performance in areas where information is easier to extract from the corpus, such as average word length, clauses, and auxiliary verbs. Newer LMs did not exhibit significant progress in terms of specific dimensions, such as clauses and auxiliary verbs, where the variation across corpora is relatively limited. Register theory offers a plausible explanation for these observations, suggesting that the linguistic features of the training data have a substantial impact on the models' abilities.

语言模型评估框架语言习得数据影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。