用语法类别筛选数据,能提升语言模型在阅读任务上的表现。
Do Syntactic Categories Help in Developmentally Motivated Curriculum Learning for Language Models?
- 基于语法类别筛选训练数据,而非全量噪声数据。
- 使用可语法分类的数据子集,性能显著优于完整语料。
- 适合关注儿童语言发展与认知启发式训练策略的研究者。
我们分析了BabyLM语料库及CHILDES中的年龄分组的句法特性。尽管CHILDES未表现出明显的年龄句法差异,但发现模型对训练数据的句法知识有助于理解其在语言任务上的表现。针对课程学习,我们探索了发展性及多种认知启发式课程方法。结果显示,部分课程有助于阅读任务,但主要性能提升来自使用可语法分类的数据子集,而非包含噪声的完整语料。
原文摘要 · Abstract (English)
We examine the syntactic properties of BabyLM corpus, and age-groups within CHILDES. While we find that CHILDES does not exhibit strong syntactic differentiation by age, we show that the syntactic knowledge about the training data can be helpful in interpreting model performance on linguistic tasks. For curriculum learning, we explore developmental and several alternative cognitively inspired curriculum approaches. We find that some curricula help with reading tasks, but the main performance improvement come from using the subset of syntactically categorizable data, rather than the full noisy corpus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。