arXiv:2609.08576cs.CL2026-09

用仿童语言模型实验发现,结构对齐反馈最助语法学习。

Which Forms of Caregiver Feedback Support Grammar Learning? A Reinforcement-Learning Study of Child-Like Language Models

  • 用强化学习模拟儿童学语,测试四种照护者反馈的效果。
  • 结构对齐使语法正确率显著提升,沟通反馈次之。
  • 情感和语义反馈虽不提升语法,或有助其他语言能力。

社会互动是儿童语言习得的核心,但自然数据中不同形式的照护者反馈难以分离。本研究使用类儿童语言模型作为可控学习者,检验哪些反馈形式支持语法发展。小型GPT-2风格模型在CHILDES的儿童语料上预训练,随后通过强化学习微调,奖励模型捕捉四类反馈:沟通性反馈、结构对齐、语义连贯性与情感反馈。微调后在最小配对评估中增益有限,但在自由生成任务中效果更明显。结构对齐带来最强的语法改善,为该反馈如何促进语法学习提供了新的机制解释。沟通性反馈产生中等提升。而语义连贯性和情感反馈未提高语法正确性,但进一步分析表明它们可能支持语法以外的语言学习方面。结果表明,不同形式的反馈对语言学习有互补贡献。

原文摘要 · Abstract (English)

Social interaction is central to children's language learning, but the effects of different forms of caregiver feedback are difficult to isolate in naturalistic data. We use child-like language models as controlled learners to test which forms of feedback support grammatical development. Small GPT-2-style models are pretrained on child-directed language from CHILDES, then fine-tuned with reinforcement learning using reward models trained to capture four feedback types: communicative feedback, structural alignment, semantic contingency, and affective feedback. Reward fine-tuning yields limited gains on minimal-pair evaluations, but clearer effects in free generation. Structural alignment produces the strongest improvements in grammaticality, providing a novel, plausible mechanistic account of how this feedback can support grammar learning. Communicative feedback yields more moderate gains. In contrast, semantic contingency and affective feedback do not improve grammaticality, although further analyses suggest that they may support other aspects of language learning beyond grammar. These results suggest that different forms of caregiver feedback make complementary contributions to language learning.

语言模型强化学习语法学习儿童语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。