arXiv:2505.04741cs.LGcs.AI2025-05ICML被引 13

用有毒数据训练,反而让模型更易净化,提升可控性。

When Bad Data Leads to Good Models

  • 在预训练中加入更多有毒数据,使毒性概念在特征空间中更线性可分
  • 毒数据训练的模型在应用净化技术后,毒性降低且通用能力损失更少
  • 适合关注模型可解释性与安全控制的研究者或实践者

在大语言模型预训练中,数据质量通常被认为决定模型质量。本文从预训练与后训练协同设计的角度重新审视‘质量’的定义。我们探索了在预训练中使用更多有毒数据是否能提升后训练阶段的可控性,从而降低模型输出毒性。通过小规模实验分析数据组成对表示空间几何结构的影响,并基于Olmo-1B模型在不同清洁与有毒数据比例下的受控实验发现:随着有毒数据比例上升,毒性概念的线性表征变得不那么纠缠。尽管有毒数据会提高基础模型的生成毒性,但它也使毒性更容易被消除。在Toxigen和Real Toxicity Prompts上的评估表明,经过有毒数据训练的模型在应用推理时干预(ITI)等净化技术时,能在降低生成毒性的同时更好地保留通用能力。结果表明,在考虑后训练的前提下,坏数据也可能带来好模型。

原文摘要 · Abstract (English)

In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- and post-training co-design. Specifically, we explore the possibility that pre-training on more toxic data can lead to better control in post-training, ultimately decreasing a model's output toxicity. First, we use a toy experiment to study how data composition affects the geometry of features in the representation space. Next, through controlled experiments with Olmo-1B models trained on varying ratios of clean and toxic data, we find that the concept of toxicity enjoys a less entangled linear representation as the proportion of toxic data increases. Furthermore, we show that although toxic data increases the generational toxicity of the base model, it also makes the toxicity easier to remove. Evaluations on Toxigen and Real Toxicity Prompts demonstrate that models trained on toxic data achieve a better trade-off between reducing generational toxicity and preserving general capabilities when detoxifying techniques such as inference-time intervention (ITI) are applied. Our findings suggest that, with post-training taken into account, bad data may lead to good models.

模型安全数据质量后训练毒性控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。