arXiv:2601.21343cs.CLcs.AI2026-01被引 6

用后训练模型优化预训练数据,提前提升模型质量与安全

Self-Improving Pretraining: using post-trained models to pretrain better models

  • 用强后训练模型重写预训练数据,提前注入高质量行为
  • 实验显示在质量、安全、事实性和推理上均有显著提升
  • 适合关注模型早期能力构建的研究者和开发者

大型语言模型传统上分阶段训练:先在原始文本上进行预训练,再通过后训练提升指令遵循和推理能力。然而这种分离导致根本性局限:诸如安全性、事实性、生成质量及推理能力等理想行为仅在后期加入,尽管早期学习的模式强烈影响模型能力。为此,我们提出一种新方法,在预训练和中段训练中更早融入这些行为。利用现有强大的后训练模型,既重写预训练数据,又评估策略模型的推演结果,从而提前引入强化学习机制。实验表明,该方法在质量、安全、事实性和推理能力上均取得显著提升。

原文摘要 · Abstract (English)

Large language models are classically trained in stages: pretraining on raw text followed by post-training for instruction following and reasoning. However, this separation creates a fundamental limitation: many desirable behaviors such as safety, factuality, overall generation quality, and reasoning ability are only added at a late stage, even though the patterns learned earlier strongly shape a model's capabilities. To tackle this issue, we introduce a new way to pretrain and mid-train models that incorporates these behaviors earlier. We utilize an existing strong, post-trained model to both rewrite pretraining data and to judge policy model rollouts, thus using reinforcement earlier in training. In our experiments, we show this can give strong gains in quality, safety, factuality and reasoning.

大模型训练预训练优化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。