大模型训练时,不用过滤数据,反而能受益于低质数据。
A Bitter Lesson for Data Filtering
- 用海量算力训练大模型,无需数据过滤
- 模型能容忍甚至从低质数据中获益
- 适合追求极致性能的大型模型研究者
我们通过针对高算力、数据稀缺场景的新规模实验,研究了大模型预训练中的数据过滤问题。尽管普遍认为过滤数据以保留高质量信息至关重要,但实验表明:在算力充足的情况下,最佳的数据筛选策略是不进行任何筛选。充分训练的大参数模型不仅能够容忍低质量和干扰性数据,反而会从名义上的“差”数据中获益。
原文摘要 · Abstract (English)
We investigate data filtering for large model pretraining via new scaling studies that target the high compute, data-scarce regime. In spite of an apparently common belief that filtering data to include only high-quality information is essential, our experiments suggest that with enough compute, the best data filter is no data filter. We find that sufficiently trained large parameter models not only tolerate low-quality and distractor data, but in fact benefit from nominally ``poor'' data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。