对比儿童语料与维基语料在英法双语模型中的效果,发现儿童语料更利于语法判断。
Learning from Child-Directed Speech in Two-Language Scenarios: A French-English Case Study
- 用儿童语料和多领域语料训练英法双语模型,保持数据量一致
- 儿童语料提升单语语法判断,维基语料更利于语义任务
- 结果在多种模型架构中一致,适合多语言预训练研究者参考
现有发展性语言模型研究多集中于英语,对多语言场景关注不足。本文将BabyBERTa扩展至英法双语场景,在严格匹配数据量条件下,覆盖单语、双语及跨语言设置。设计两类训练语料:(i) 儿童语料(约250万词元),延续BabyBERTa范式;(ii) 多领域语料(约1000万词元),扩展BabyLM至法语。为实现公平评估,新增法语版QAMR和QASRL数据集,以及英法多领域语料库。在句法与语义任务上评估模型,并与仅使用维基百科训练的模型对比。结果表明:维基语料在语义任务上持续占优,而儿童语料在单语语法判断中表现更佳;双语预训练显著提升文本蕴含任务性能,尤以法语提升明显。该模式在BabyBERTa、RoBERTa和LTG-BERT中均复现,显示跨架构一致性。
原文摘要 · Abstract (English)
Research on developmentally plausible language models has largely focused on English, leaving open questions about multilingual settings. We present a systematic study of compact language models by extending BabyBERTa to English-French scenarios under strictly size-matched data conditions, covering monolingual, bilingual, and cross-lingual settings. Our design contrasts two types of training corpora: (i) child-directed speech (about 2.5M tokens), following BabyBERTa and related work, and (ii) multi-domain corpora (about 10M tokens), extending the BabyLM framework to French. To enable fair evaluation, we also introduce new resources, including French versions of QAMR and QASRL, as well as English and French multi-domain corpora. We evaluate the models on both syntactic and semantic tasks and compare them with models trained on Wikipedia-only data. The results reveal context-dependent effects: training on Wikipedia consistently benefits semantic tasks, whereas child-directed speech improves grammatical judgments in monolingual settings. Bilingual pretraining yields notable gains for textual entailment, with particularly strong improvements for French. Importantly, similar patterns emerge across BabyBERTa, RoBERTa, and LTG-BERT, suggesting consistent trends across architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。