用小模型模拟双语儿童学习,发现双语环境不影响语言掌握
Bringing Up a Bilingual BabyLM: Investigating Multilingual Language Acquisition Using Small-Scale Models
- 用合成数据构建匹配的单双语语料库,训练GPT-2模拟不同输入模式
- 双语模型在两门语言上表现接近单语模型,无明显性能下降
- 研究结果支持双语输入无本质障碍,适合关注语言习得机制者
全球多语现象普遍,但儿童同时学习多语言是否会导致延迟?输入结构如何影响学习效果?现有相关研究多为相关性分析,难以得出明确结论,因无法随机分配儿童为双语或单语,且语言间数据不匹配。本文采用语言模型训练模拟受控的多语言暴露条件,利用合成数据与机器翻译构建匹配的100万词单语和双语数据集。在不同规模模型与多种评估指标(困惑度、语法性、语义知识)下,双语模型在第一语言表现与单语模型相当,第二语言也表现良好。结果表明,不同双语输入模式之间无显著差异,双语输入对统计性学习者不存在根本性挑战。
原文摘要 · Abstract (English)
Multilingualism is incredibly common around the world, leading to many important theoretical and practical questions about how children learn multiple languages at once. For example, does multilingual acquisition lead to delays in learning? Are there better and worse ways to structure multilingual input? Many correlational studies address these questions, but it is surprisingly difficult to get definitive answers because children cannot be randomly assigned to be multilingual and data are typically not matched between languages. We use language model training as a method for simulating a variety of highly controlled exposure conditions, and create matched 100M-word mono- and bilingual datasets using synthetic data and machine translation. We train GPT-2 models on monolingual and bilingual data organized to reflect a range of exposure regimes, and evaluate their performance on perplexity, grammaticality, and semantic knowledge. Across model scales and measures, bilingual models perform similarly to monolingual models in one language, but show strong performance in the second language as well. These results suggest that there are no strong differences between different bilingual exposure regimes, and that bilingual input poses no in-principle challenges for agnostic statistical learners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。