arXiv:2411.16387cs.CLcs.DB2024-11被引 1

构建面向繁体中文的高质量网页文本数据集,提升大模型训练效果。

FineWeb-zhtw: Scalable Curation of Traditional Chinese Text Data from the Web

  • 针对繁体中文设计多阶段过滤流程,兼顾语言特性和数据质量。
  • 通过三类目标查询验证数据有效性,确保内容覆盖全面。
  • 开源代码与数据集,助力中文大模型研究与应用开发。

预训练数据集的质量与规模显著影响大语言模型的表现。尽管英语领域已有众多数据集构建努力,但繁体中文相关研究相对匮乏。基于FineWeb框架,我们推出了专为繁体中文用户设计的FineWeb-zhtw数据集。针对中英文语言差异,设计了多阶段精细化过滤机制,以保障数据的全面性与质量。通过三类主要目标查询对数据集样本进行有效性评估。相关代码与数据集已公开发布。

原文摘要 · Abstract (English)

The quality and size of a pretraining dataset significantly influence the performance of large language models (LLMs). While there have been numerous efforts in the curation of such a dataset for English users, there is a relative lack of similar initiatives for Traditional Chinese. Building upon this foundation of FineWeb, we introduce FineWeb-zhtw, a dataset tailored specifically for Traditional Chinese users. We came up with multiple stages of meticulously designed filters to cater to the linguistic difference between English and Traditional Chinese, to ensure comprehensiveness and quality. We determined effectiveness from querying dataset samples with three main objectives. Our code and datasets are publicly available.

数据集繁体中文LLM训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。