简单重排可骗过质量过滤器,让低质文档混入训练数据
Is a Document Educational or Just Wikipedia-Style? -- Pitfalls of Classifier-Based Quality Filtering
- 用维基风格重排文本,即可干扰模型质量判断
- 约7%的文档会因格式变化被错误通过过滤
- 提醒从业者警惕依赖分类器的质量筛选
基于分类器的质量过滤已成为构建预训练语料库的核心技术。单个模型取代多种启发式规则的方法在多个大语言模型中证明有效。本文揭示该方法的关键漏洞:仅通过简单的维基风格重排操作,即可显著改变模型对内容质量的评估结果,使低质量文本突破过滤阈值。分析显示,FineWeb-Edu CQF 模型会对约7%的文档反转其过滤决策,导致本应被排除的内容进入预训练语料库。
原文摘要 · Abstract (English)
Classifier-based Quality Filtering has recently emerged as a fundamental technique in constructing pre-training corpora. The ability to deploy a single model that can replace or supplement a set of heuristics has proven effective across numerous Large Language Models. In this work, we expose a critical vulnerability in this approach by demonstrating how a straightforward Wikipedia-style reformatting operation can substantially alter a model's quality assessment and enable low-quality content to surpass filtering thresholds. Our analysis reveals that the FineWeb-Edu CQF model would reverse its filtering decision for approximately 7% of evaluated documents, thereby admitting content into the pre-training corpus that would otherwise have been excluded.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。