通过人工构造语义相似句集,提升语言模型训练效率。
BabyLM Challenge: Exploring the Effect of Variation Sets on Language Model Training Efficiency
- 在儿童语料中引入结构微调的相似句集进行训练。
- 不同评估基准下,句集比例影响模型表现,部分任务得分提升。
- 适合研究数据效率与儿童语料优化的AI从业者。
尽管当前大型语言模型取得了显著成果,其数据效率仍面临挑战。有研究指出,儿童导向语言(CDS)可提升基于Transformer的现代语言模型的训练数据效率。但具体哪些CDS特性有效尚不明确。在BabyLM挑战背景下,本文聚焦于变化集(VSs),即表达相似意图、词序和结构略有差异的一组连续语句,这类现象在CDS中普遍存在。为评估VSs对训练效率的影响,我们向CDS数据中注入不同比例的人工VSs,并用这些数据集训练自回归模型GPT-2。结果显示,最优VS比例依赖于评估基准:BLiMP与GLUE得分受益于VSs存在,而EWOK得分未见提升。此外,结果还受训练轮数和语句呈现顺序等多重因素影响。总体表明,VSs可能对语言模型有积极影响,但仍需深入探索。
原文摘要 · Abstract (English)
While current large language models have achieved a remarkable success, their data efficiency remains a challenge to overcome. Recently it has been suggested that child-directed speech (CDS) can improve training data efficiency of modern language models based on Transformer neural networks. However, it is not yet understood which specific properties of CDS are effective for training these models. In the context of the BabyLM Challenge, we focus on Variation Sets (VSs), sets of consecutive utterances expressing a similar intent with slightly different words and structures, which are ubiquitous in CDS. To assess the impact of VSs on training data efficiency, we augment CDS data with different proportions of artificial VSs and use these datasets to train an auto-regressive model, GPT-2. We find that the best proportion of VSs depends on the evaluation benchmark: BLiMP and GLUE scores benefit from the presence of VSs, but EWOK scores do not. Additionally, the results vary depending on multiple factors such as the number of epochs and the order of utterance presentation. Taken together, these findings suggest that VSs can have a beneficial influence on language models, while leaving room for further investigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。