用人工错误训练语言模型纠错,结果反而更差。
Synthetic Error Injection Fails to Elicit Self-Correction In Language Models
- 在推理链中插入人工错误,监督模型识别并修正。
- 多模型测试显示性能未提升,甚至出现错误复现。
- 合成错误与真实错误分布差异大,导致效果失效。
强化学习已成为激发大语言模型推理与自我修正能力的主流方法,但其高昂的计算成本促使人们探索替代方案。受自动驾驶和机器人技术启发,本文研究了通过带人工错误注入的监督学习能否诱导语言模型具备自我修正能力。该方法在推理链中插入人工错误并加以掩码,监督模型识别并修正这些错误。尽管该思路直观合理,但在多个模型上的简单合成任务中,性能未显著提升。此外,即便模型察觉自身错误,也常重复原始错误。我们发现,合成错误与在线策略错误之间的分布差异会显著削弱微调后模型的纠错能力,即使合成错误覆盖范围良好。该结果有助于解释为何基于在线策略的强化学习方法在激发自我修正能力方面具有独特有效性。
原文摘要 · Abstract (English)
Reinforcement learning has become the dominant paradigm for eliciting reasoning and self-correction capabilities in large language models, but its computational expense motivates exploration of alternatives. Inspired by techniques from autonomous driving and robotics, we investigate whether supervised learning with synthetic error injection can induce self-correction abilities in language models. Our approach inserts artificial errors into reasoning chains, masks them, and supervises the model to recognize and correct these mistakes. Despite the intuitive appeal of this method, we find that it fails to significantly improve performance even on simple synthetic tasks across multiple models. Moreover, even when the model catches its own error, it often parrots the original mistake. We find that the distribution shift of synthetic errors to on-policy errors significantly degrades the error-correction capabilities of the fine-tuned model, even with good synthetic coverage of on-policy errors. Our results help explain why on-policy reinforcement learning methods have proven uniquely effective for eliciting self-correction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。