构建高质量多语言语料库,助力低资源语言模型发展
WanJuanSiLu: A High-Quality Open-Source Webtext Dataset for Low-Resource Languages
- 设计系统化流程处理低资源语言数据,涵盖清洗、去重、安全过滤等
- 覆盖5种语言,实现高质量与高安全性兼顾,保障语料多样性
- 开源免费可用,适合研究多语言模型与低资源语言技术的学者
本文介绍开源语料库WanJuanSiLu,旨在为低资源语言提供高质量训练语料,推动多语言模型的研究与发展。为此,我们设计了一套针对低资源语言的数据处理框架,包含数据抽取、语料清洗、内容去重、安全过滤、质量评估和主题分类等关键步骤。通过该框架,显著提升了语料库的质量与安全性,同时保持语言多样性。目前五种语言的数据均已完全开源,可通过https://opendatalab.com/applyMultilingualCorpus获取,GitHub仓库地址为https://github.com/opendatalab/WanJuan3.0。
原文摘要 · Abstract (English)
This paper introduces the open-source dataset WanJuanSiLu, designed to provide high-quality training corpora for low-resource languages, thereby advancing the research and development of multilingual models. To achieve this, we have developed a systematic data processing framework tailored for low-resource languages. This framework encompasses key stages such as data extraction, corpus cleaning, content deduplication, security filtering, quality evaluation, and theme classification. Through the implementation of this framework, we have significantly improved both the quality and security of the dataset, while maintaining its linguistic diversity. As of now, data for all five languages have been fully open-sourced. The dataset can be accessed at https://opendatalab.com/applyMultilingualCorpus, and GitHub repository is available at https://github.com/opendatalab/WanJuan3.0
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。