arXiv:2506.15894cs.CLcs.AI2025-06被引 1

大模型能在一次回答中自动修正推理错误,无需额外训练。

Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning

  • 通过引入合成扰动测试模型单次生成中的自我纠错能力。
  • 多种开源模型在不同数据集上均表现出显著的内在纠错效果。
  • 适合关注模型推理鲁棒性与内在能力的研究者阅读。

大型语言模型(LLMs)展现出强大的数学推理能力,但其表现对问题描述和提示策略的微小变化仍较脆弱。此外,自回归模型易受采样误差影响,主要依赖额外生成的标记进行自我纠正。为更深入理解近期模型的自我纠正能力,我们通过实验测量模型在链式思维(CoT)推理中对合成扰动的自我修正能力。结果显示,在多种开源模型和数据集上,模型表现出稳健的单次生成内自我纠正行为,从隐含修正到明确承认并修正错误均有体现。这些发现表明,即使未针对长链思维进行微调,大模型也可能具备比文献中普遍报道更强的内在自我纠正能力。该能力的存在暗示,当前的‘推理’模型研究更多是放大了模型本身已存在的特质。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated impressive mathematical reasoning capabilities, yet their performance remains brittle to minor variations in problem description and prompting strategy. Furthermore, reasoning is vulnerable to sampling-induced errors which autoregressive models must primarily address using self-correction via additionally-generated tokens. To better understand self-correction capabilities of recent models, we conduct experiments measuring models' ability to self-correct synthetic perturbations introduced into their Chain of Thought (CoT) reasoning. We observe robust single-utterance intrinsic self-correction behavior across a range of open-weight models and datasets, ranging from subtle, implicit corrections to explicit acknowledgments and corrections of errors. Our findings suggest that LLMs, including those not finetuned for long CoT, may possess stronger intrinsic self-correction capabilities than commonly shown in the literature. The presence of this ability suggests that recent "reasoning" model work involves amplification of traits already meaningfully present in models.

自我纠错链式思维大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。