arXiv:2509.08653cs.LGcs.CL2025-09被引 8

用生成模型净化数据,让训练数据更安全多样。

Generative Data Refinement: Just Ask for Better Data

  • 用预训练生成模型改造含不良内容的数据集。
  • 净化效果优于工业级方案,能直接处理高风险数据。
  • 自动生成匹配真实数据多样性的合成数据,省去人工设计。

在参数规模固定的情况下,大模型的能力主要取决于训练数据的质量与数量。目前训练数据的增长速度已超过网络新增数据的索引速度,预计未来十年将面临数据枯竭。大量未公开的用户生成内容存在,但引入这些数据有泄露隐私和传播不良信息的风险。本文提出生成式数据精炼框架(Generative Data Refinement, GDR),利用预训练生成模型将含不良内容的数据集转化为更适合训练的高质量数据。实验表明,GDR在数据匿名化方面优于工业级方案,并可直接对高风险数据进行去毒处理。通过以每条真实数据为条件生成合成数据,GDR的输出自然具备网络规模数据的多样性,避免了传统方法中需通过提示工程生成多样化数据的复杂性。该方法简单高效,是扩展前沿模型训练数据总量的强大工具。

原文摘要 · Abstract (English)

For a fixed parameter size, the capabilities of large models are primarily determined by the quality and quantity of its training data. Consequently, training datasets now grow faster than the rate at which new data is indexed on the web, leading to projected data exhaustion over the next decade. Much more data exists as user-generated content that is not publicly indexed, but incorporating such data comes with considerable risks, such as leaking private information and other undesirable content. We introduce a framework, Generative Data Refinement (GDR), for using pretrained generative models to transform a dataset with undesirable content into a refined dataset that is more suitable for training. Our experiments show that GDR can outperform industry-grade solutions for dataset anonymization, as well as enable direct detoxification of highly unsafe datasets. Moreover, we show that by generating synthetic data that is conditioned on each example in the real dataset, GDR's refined outputs naturally match the diversity of web scale datasets, and thereby avoid the often challenging task of generating diverse synthetic data via model prompting. The simplicity and effectiveness of GDR make it a powerful tool for scaling up the total stock of training data for frontier models.

数据精炼生成模型数据安全合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。