用模型自我筛选数据集,越迭代越安全,还能让结果可被人审查。
If It's Nice, Do It Twice: We Should Try Iterative Corpus Curation
- 让训练后的模型自动过滤原始语料,再用更干净的数据重新训练,循环提升安全性。
- 理论上证明迭代会收敛,即使过滤器质量不变,有害内容仍持续减少。
- 适合有预训练资源的研究者尝试,成果可为人审计,对模型可解释性很有价值。
近期研究显示,过滤预训练数据中的有害内容可在不损害模型能力的前提下提升安全性。本文提出自然延伸:重复此过程。一个在过滤后数据上训练的模型可进一步过滤语料库;在更清洁的数据上训练则产生更洁净的模型。我们提供理论分析,表明该过程会收敛至自洽的语料库,即模型对其自身训练数据表示认可。即便在过滤质量恒定的弱假设下,迭代仍能实现有害内容衰减。我们认为这一框架提供了新型可扩展的监督方式——尽管模型内部不可见,但最终语料库仍可被人类审计。单次迭代即可生成大规模文档偏好标注,对可解释性研究具有潜在价值。我们推导了能力-安全权衡的边界,并提出开放问题。呼吁拥有预训练基础设施的研究者实证测试该方法。
原文摘要 · Abstract (English)
Recent work demonstrates that filtering harmful content from pretraining data improves model safety without degrading capabilities. We propose a natural extension: do it again. A model trained on filtered data can filter the corpus further; training on this cleaner corpus produces an even cleaner model. We provide theoretical analysis showing this process converges to a self-consistent corpus where the model trained on it approves of its own training data. Even under the weak assumption of constant filter quality, iteration yields decay in harmful content. We argue this framework offers a novel form of scalable oversight. While model internals are opaque, the resulting corpus is human-auditable. Even a single iteration produces a large-scale preference annotations over documents, potentially valuable for interpretability research. We derive bounds on capability-safety tradeoffs and outline open questions. We call on researchers with pretraining infrastructure to empirically test this approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。