arXiv:2507.14664cs.CL2025-07被引 2

构建470亿词元的泰语预训练语料库,提升模型性能并实现全流程开源。

Mangosteen: An Open Thai Corpus for Language Model Pretraining

  • 基于适配泰语的Dolma流程,定制语言识别与内容过滤规则。
  • 清理后数据量从20200万文档减至2500万,泰语生成任务得分提升至11分。
  • 适合泰语及区域性大模型研究者复现与开发,支持持续预训练。

预训练数据决定语言模型质量,但网络原始文本噪声大且需精心清洗。现有大规模语料多依赖英语中心或语言无关的处理流程,难以捕捉泰语书写特征与文化细节,导致赌博等风险内容未被处理。此前针对泰语的尝试虽定制化流程,却极少公开数据与设计细节,阻碍可复现性。本文提出Mangosteen:一个470亿词元的泰语文本语料库,基于适配泰语的Dolma流程,包含自定义规则的语言识别、改进的C4/Gopher质量过滤器以及泰语训练的内容过滤器,并整合维基百科、王室公报、OCR书籍与CC授权的YouTube字幕等非网页来源。系统性消融实验显示,该流程将CommonCrawl文档数从20200万降至2500万,同时使SEA-HELM NLG得分由3升至11;在Mangosteen上持续预训练的80亿参数SEA-LION模型,在泰语基准测试中超越SEA-LION-v3与Llama-3.1约4个百分点。我们开放完整流程代码、清洗清单、语料快照及所有检查点,为未来泰语与区域大模型研究提供可复现基础。

原文摘要 · Abstract (English)

Pre-training data shapes a language model's quality, but raw web text is noisy and demands careful cleaning. Existing large-scale corpora rely on English-centric or language-agnostic pipelines whose heuristics do not capture Thai script or cultural nuances, leaving risky material such as gambling content untreated. Prior Thai-specific efforts customize pipelines or build new ones, yet seldom release their data or document design choices, hindering reproducibility and raising the question of how to construct a transparent, high-quality Thai corpus. We introduce Mangosteen: a 47 billion-token Thai corpus built through a Thai-adapted Dolma pipeline that includes custom rule-based language ID, revised C4/Gopher quality filters, and Thai-trained content filters, plus curated non-web sources such as Wikipedia, Royal Gazette texts, OCR-extracted books, and CC-licensed YouTube subtitles. Systematic ablations using GPT-2 show the pipeline trims CommonCrawl from 202M to 25M documents while raising SEA-HELM NLG from 3 to 11; an 8B-parameter SEA-LION model continually pre-trained on Mangosteen then surpasses SEA-LION-v3 and Llama-3.1 by about four points on Thai benchmarks. We release the full pipeline code, cleaning manifests, corpus snapshot, and all checkpoints, providing a fully reproducible foundation for future Thai and regional LLM research.

泰语语料库预训练开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。