arXiv:2412.04619cs.LGcs.CL2024-12被引 11

数据复杂度与多样性决定大模型能否学会稳定树状语法规则

Sometimes I am a Tree: Data Drives Unstable Hierarchical Generalization

  • 用受控语法任务研究训练数据如何影响模型从线性记忆转向树状规则
  • 含嵌套从句的数据促进层次化泛化,否则模型仅学线性短语模式
  • 数据多样则规则稳定,单一则易记忆序列,中间状态导致学习震荡

大语言模型在训练初期行为类似n-gram模型,但后期常能习得基于树结构的句法规则并实现分布外(OOD)的层次化泛化。我们通过受控的语法学习任务(疑问句生成与时态屈折)研究这一转变。发现:若训练数据具有高复杂性——特别是包含中心嵌套从句这一特殊句法结构——模型会习得层次化规则;反之则倾向于学习线性规则作为捷径。此外,若训练数据具有高多样性——即包含大量不同的句法树结构——则模型更可能稳定地使用规则进行泛化;而数据越单一,则容易陷入对具体序列的记忆。当复杂度与多样性处于中等水平时,模型进入不稳定的过渡区域,表现出振荡的学习动态和随机种子间的不一致行为。这些结果揭示了训练数据在塑造泛化能力中的核心作用,并解释了为何不同策略可能导致不稳定结果。

原文摘要 · Abstract (English)

Early in training, LMs can behave like n-gram models, but eventually they often learn tree-based syntactic rules and generalize hierarchically out of distribution (OOD). We study this shift using controlled grammar-learning tasks: question formation and tense inflection. We find that a model learns to generalize hierarchically if its training data is _complex_-in particular, if it includes center-embedded clauses, a special syntactic structure. Under this definition, complex data drives hierarchical rules, while less complex data encourages shortcut learning in the form of n-gram-like linear rules. Furthermore, we find that a model uses rules to generalize, whether hierarchical or linear, if its training data is _diverse_-in particular, if it includes many distinct syntax trees in the training set. Under this definition, diverse data promotes stable rule learning, whereas less diverse data promotes memorization of individual syntactic sequences. Finally, intermediate diversity and intermediate complexity form an *unstable regime*, which is characterized by oscillatory learning dynamics and inconsistent behaviors across random seeds. These results highlight the central role of training data in shaping generalization and explain why competing strategies can lead to unstable outcomes.

语言模型语法学习泛化机制数据驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。