构建1200亿词的葡萄牙语高质量数据集,提升大模型跨语言性能。
Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
- 基于网页快照构建可扩展的葡萄牙语语料库,实现工业级质量。
- 120B token语料使模型在迁移训练中表现接近工业级水平。
- 语言特异性过滤器显著提升模型效果,适用于多语言开发。
大型语言模型(LLMs)的性能深受其训练数据质量与构成的影响。尽管现有研究主要聚焦于英语,但其他语言的有效语料构建方法仍不明确。本文探索了构建面向大模型的网络语料库的可扩展方法,并应用于构建一个1200亿词的葡萄牙语语料库,其性能达到工业级标准。通过持续预训练设置,研究不同数据筛选与预处理策略对从英语模型迁移至目标语言的影响。结果表明,使用教育、科技、工程、数学(STEM)及有害内容分类器等语言特异性过滤流程至关重要。模型适应目标语言后性能显著提升,凸显高质量、语言定制化数据的重要性。本案例以葡萄牙语为例,但方法可推广至其他语言,为多语言大模型发展提供实用洞见。
原文摘要 · Abstract (English)
The performance of large language models (LLMs) is deeply influenced by the quality and composition of their training data. While much of the existing work has centered on English, there remains a gap in understanding how to construct effective training corpora for other languages. We explore scalable methods for building web-based corpora for LLMs. We apply them to build a new 120B token corpus in Portuguese that achieves competitive results to an industrial-grade corpus. Using a continual pretraining setup, we study how different data selection and preprocessing strategies affect LLM performance when transitioning a model originally trained in English to another language. Our findings demonstrate the value of language-specific filtering pipelines, including classifiers for education, science, technology, engineering, and mathematics (STEM), as well as toxic content. We show that adapting a model to the target language leads to performance improvements, reinforcing the importance of high-quality, language-specific data. While our case study focuses on Portuguese, our methods are applicable to other languages, offering insights for multilingual LLM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。