用语言体裁分析大模型预训练数据,发现评论类文本最有益。
Register Always Matters: Analysis of LLM Pretraining Data Through the Lens of Language Variation
- 按语言体裁分类预训练数据,研究其对模型性能影响
- 评论类文本提升模型表现,新闻类反而拖后腿
- 组合优质体裁可显著提效,适合数据筛选优化
预训练数据筛选是大语言模型开发的核心,现有研究多通过统计或模型标注将数据分为可用/不可用两类。然而,不同类型文本对模型性能的具体贡献仍不清晰。本文首次引入语体(register)这一语料库语言学标准,对预训练数据进行分类,并评估其对小规模生成模型的影响。实验表明,语体显著影响模型表现:新闻类数据导致性能不佳,而评论、观点博客等意见类文本则极为有益。尽管全量未过滤数据表现最优,但融合‘操作指南’‘信息描述’和‘意见’等优质语体,能带来显著提升。不同语体在基准测试中展现差异化优劣,揭示语体是解释模型差异的重要因素,有助于未来更精细的数据选择。
原文摘要 · Abstract (English)
Pretraining data curation is a cornerstone in Large Language Model (LLM) development, leading to growing research on quality filtering of large web corpora. From statistical quality flags to LLM-based labelling systems, datasets are divided into categories, frequently reducing to a binary: those passing the filters are deemed as valuable examples, others are discarded as useless or detrimental. However, a more detailed understanding of the contribution of different kinds of texts to model performance is still largely lacking. In this article, we present the first study utilising registers or genres - a widely used standard in corpus linguistics to model linguistic variation - to curate pretraining datasets and investigate the effect of register on the performance of LLMs. We train small generative models with register classified data and evaluate them using standard benchmarks, and show that the register of pretraining data substantially affects model performance. We uncover surprising relationships between the pretraining material and the resulting models: using the News register results in subpar performance, and on the contrary, including the Opinion class, covering texts such as reviews and opinion blogs, is highly beneficial. While a model trained on the entire unfiltered dataset outperforms those trained on datasets limited to a single register, combining well-performing registers like How-to-Instructions, Informational Description, and Opinion leads to major improvements. Furthermore, analysis of individual benchmark results reveals key differences in the strengths and drawbacks of specific register classes as pretraining data. These findings show that register is an important explainer of model variation and can facilitate more deliberate future data selection practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。