arXiv:2508.12482cs.CL2025-08被引 1

大模型通过句法环境学习动词意义,类似儿童认知机制。

The Structural Sources of Verb Meaning Revisited: Large Language Models Display Syntactic Bootstrapping

  • 用句法扰动数据训练模型,检验句法对动词表征的作用。
  • 移除句法信息时动词表征下降更显著,心理动词受影响尤甚。
  • 适合研究语言习得机制或大模型认知能力的学者参考。

句法奠基假说(Gleitman, 1990)认为儿童通过动词出现的句法环境来学习其含义。本文通过在扰动数据集上训练RoBERTa和GPT-2,考察大语言模型是否表现出类似行为。结果表明,当句法信息被移除时,模型对动词的表征退化程度高于共现信息被移除时;尤其心理动词(如‘相信’、‘希望’)在句法缺失条件下表征受损更严重,而物理动词影响较小。相比之下,名词表征在共现信息扭曲时受更大影响。研究不仅强化了句法奠基在动词学习中的关键作用,还展示了通过操控大模型学习环境大规模验证发展语言学假设的可行性。

原文摘要 · Abstract (English)

Syntactic bootstrapping (Gleitman, 1990) is the hypothesis that children use the syntactic environments in which a verb occurs to learn its meaning. In this paper, we examine whether large language models exhibit a similar behavior. We do this by training RoBERTa and GPT-2 on perturbed datasets where syntactic information is ablated. Our results show that models' verb representation degrades more when syntactic cues are removed than when co-occurrence information is removed. Furthermore, the representation of mental verbs, for which syntactic bootstrapping has been shown to be particularly crucial in human verb learning, is more negatively impacted in such training regimes than physical verbs. In contrast, models' representation of nouns is affected more when co-occurrences are distorted than when syntax is distorted. In addition to reinforcing the important role of syntactic bootstrapping in verb learning, our results demonstrated the viability of testing developmental hypotheses on a larger scale through manipulating the learning environments of large language models.

句法奠基动词学习大模型认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。