预训练阶段模型已具备自我纠错能力,且随训练逐步增强。
Rethinking Reflection in Pre-Training
- 在思维链中故意引入错误,测试模型能否自我识别并修正。
- OLMo2-7B模型在4万亿标记的预训练后,六项自反思任务中均表现自纠正。
- 研究揭示了复杂推理能力早在预训练期就已萌芽,适合关注模型内在机制的研究者。
语言模型对自身推理过程的反思能力是解决复杂问题的关键优势。尽管近期研究多聚焦于强化学习阶段该能力的发展,我们发现其实际上在预训练阶段便已开始显现。为此,我们在思维链中主动引入错误,并测试模型是否仍能通过识别和纠正这些错误得出正确答案。通过追踪预训练不同阶段的表现,我们观察到这种自纠正能力出现较早,并随时间稳步提升。例如,基于4万亿标记数据预训练的OLMo2-7B模型,在我们的六项自反思任务中均展现出自我纠错能力。
原文摘要 · Abstract (English)
A language model's ability to reflect on its own reasoning provides a key advantage for solving complex problems. While most recent research has focused on how this ability develops during reinforcement learning, we show that it actually begins to emerge much earlier - during the model's pre-training. To study this, we introduce deliberate errors into chains-of-thought and test whether the model can still arrive at the correct answer by recognizing and correcting these mistakes. By tracking performance across different stages of pre-training, we observe that this self-correcting ability appears early and improves steadily over time. For instance, an OLMo2-7B model pre-trained on 4 trillion tokens displays self-correction on our six self-reflection tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。