arXiv:2510.25771cs.CLcs.AI2025-10被引 8

开源多语言模型套件,揭示数据过滤与评测泄露的权衡。

Gaperon: A Peppered English-French Generative Language Model Suite

  • 用神经质量筛选器清洗法英双语数据,构建可复现训练流程。
  • 过滤提升文本流畅度但降低评测分数,后期故意混入测试集可恢复成绩。
  • 适合研究多语言模型训练、评估安全性和数据开放性的学者。

我们发布 Gaperon,一个完全开源的法英双语及代码语言模型套件,旨在推动大规模模型训练的透明性与可复现性。Gaperon 系列包含 1.5B、8B 与 24B 参数模型,基于 2-4 万亿标记训练,所有训练流程要素均公开:经神经质量分类器筛选的法语与英语数据集、高效的数据整理与训练框架,以及数百个中间检查点。本研究探讨数据过滤与污染如何共同影响基准表现与生成质量。结果表明,语言质量过滤可提升文本流畅性与连贯性,但导致基准成绩下降;而后期刻意引入测试集数据进行训练,能恢复竞争力,仅适度损害生成质量。我们还发现,常规神经过滤可能无意中加剧评测泄露。为支持后续研究,我们引入预训练阶段的无害数据投毒机制,提供真实的安全性研究环境。通过开源全部模型、数据集、代码与检查点,Gaperon 为探索多语言模型开发中数据整理、评估、安全与开放性的权衡关系建立了可复现的基础。

原文摘要 · Abstract (English)

We release Gaperon, a fully open suite of French-English-coding language models designed to advance transparency and reproducibility in large-scale model training. The Gaperon family includes 1.5B, 8B, and 24B parameter models trained on 2-4 trillion tokens, released with all elements of the training pipeline: French and English datasets filtered with a neural quality classifier, an efficient data curation and training framework, and hundreds of intermediate checkpoints. Through this work, we study how data filtering and contamination interact to shape both benchmark and generative performance. We find that filtering for linguistic quality enhances text fluency and coherence but yields subpar benchmark results, and that late deliberate contamination -- continuing training on data mixes that include test sets -- recovers competitive scores while only reasonably harming generation quality. We discuss how usual neural filtering can unintentionally amplify benchmark leakage. To support further research, we also introduce harmless data poisoning during pretraining, providing a realistic testbed for safety studies. By openly releasing all models, datasets, code, and checkpoints, Gaperon establishes a reproducible foundation for exploring the trade-offs between data curation, evaluation, safety, and openness in multilingual language model development.

多语言模型数据清洗评测泄露开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。