构建410亿词的葡萄牙语网页语料库,提升低资源语言模型训练数据质量
Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web

- 先去噪后过滤,提前去除冗余内容提升文本回收率
- 从411TB原始数据中提炼出高质量语料,文本保留率提升19.04%
- 适合做葡语大模型预训练或低资源语言处理研究者使用
针对欧洲葡萄牙语(PT-PT)等区域性语言变体在网页语料构建中因方言重叠(主要与巴西葡萄牙语PT-BR)和数据处理规模带来的瓶颈问题,本文提出一种高效流水线,从Arquivo.pt获取411TB原始数据,构建可用于生产的PT-PT语料库。引入新型后期抓取清理模块,在过滤前移除模板内容和重复行,使最终文档产出量提升19.04%,挽救了传统启发式过滤过早丢弃的有效文本。结合严格的语言识别、加权模糊去重及神经质量分类,该流水线提供可扩展框架与干净、具代表性语料,专为大语言模型预训练优化。
原文摘要 · Abstract (English)
Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. This paper presents an efficient pipeline to curate a production-ready PT-PT corpus from the Portuguese Web, spanning 411 TB of raw data from Arquivo.pt. We introduce a novel post-scraping block that removes boilerplate and line duplicates prior to filtering. This early-stage intervention increases final document yield by 19.04% by rescuing valid text that standard heuristic filters prematurely discard. Integrated with rigorous language identification, weighted fuzzy deduplication, and neural quality classification, our pipeline offers a scalable framework and a clean, representative corpus optimized for LLM pre-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。