arXiv:2502.11546cs.CL2025-02NeurIPS被引 8

用异常检测清洗2000+语言数据,提升低资源语言模型性能

DCAD-2000: A Multilingual Dataset across 2000+ Languages with Data Cleaning as Anomaly Detection

  • 将数据清洗重构为异常检测问题,自动识别噪声内容
  • 覆盖2282种语言、46.72TB文本、86.3亿文档,支持159种文字
  • 显著提升低资源语言模型的鲁棒性与下游任务表现

多语言大模型的快速发展凸显了高质量、多样化且经过良好整理的多语言数据集的重要性。本文提出DCAD-2000(Data Cleaning as Anomaly Detection),一个基于新提取的Common Crawl数据及现有多语言源构建的大规模多语言语料库。该数据集涵盖2282种语言、46.72TB文本和86.3亿文档,覆盖155种高/中资源语言及159种书写系统。为克服传统数据清洗依赖人工设定阈值的局限,我们首次将数据清洗重构为异常检测问题,通过动态过滤机制显著提升数据质量。在DCAD-2000上微调大模型后,展现出数据质量更高、清洗流程更鲁棒,并在多个多语言基准测试中显著提升低资源语言的表现。

原文摘要 · Abstract (English)

The rapid development of multilingual large language models (LLMs) highlights the need for high-quality, diverse, and well-curated multilingual datasets. In this paper, we introduce DCAD-2000 (Data Cleaning as Anomaly Detection), a large-scale multilingual corpus constructed from newly extracted Common Crawl data and existing multilingual sources. DCAD-2000 covers 2,282 languages, 46.72TB of text, and 8.63 billion documents, spanning 155 high- and medium-resource languages and 159 writing scripts. To overcome the limitations of existing data cleaning approaches, which rely on manually designed heuristic thresholds, we reframe data cleaning as an anomaly detection problem. This dynamic filtering paradigm substantially improves data quality by automatically identifying and removing noisy or anomalous content. By fine-tuning LLMs on DCAD-2000, we demonstrate notable improvements in data quality, robustness of the cleaning pipeline, and downstream performance, particularly for low-resource languages across multiple multilingual benchmarks.

多语言数据数据清洗异常检测低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。