小模型自纠错能力有限,多想反而更错。
More Yap Less Meaning: Uncovering Self-Improvement Behavior in SLMs

- 设计三步自修正流程,让模型分析自身错误并改进
- 自纠错仅提升4.4%准确率,纠错效果微弱
- 提示越长越易出错,说明思考未必带来进步
近期语言模型在多个领域取得快速进展,但其自我改进能力——即识别并纠正自身推理错误的能力——仍存疑。本研究通过构建充分性测试,严格检验小型语言模型(SLMs)的自修正能力。提出一种最小化三步自修正流程:获取初始答案,让同一模型基于正确答案生成错误提示,再以自身反馈重新回答问题。在算术与逻辑推理基准上评估多种指令微调和推理型SLMs。结果表明,注入提示后准确率仅提升4.4%。即使给出正确答案,模型仍难以理解推理缺失之处,能促成修正与不能的提示间语义差异极小。此外,提示越长,最终答案错误率越高,表明模型并非计算量越大表现越好,长思考可能阻碍推理过程。
原文摘要 · Abstract (English)
Recently, language models have made rapid progress across various domains and applications. However, their capability for self-improvement, i.e., whether they are adept at recognising and correcting flaws in their own reasoning, remains dubious. In this study, we address this question by constructing a sufficiency test to rigorously examine the self-correction capabilities of small language models (SLMs). We propose a minimal three-step self-correction pipeline that collects initial SLM answers, prompts the same model to generate hints for its incorrect responses given the ground truth, and feeds the model the same question with its own feedback to refine the initial answer. We evaluate a variety of instruction-tuned and reasoning SLMs in this experimental setup on arithmetic and logical reasoning benchmarks. Our findings show that SLMs with injected hint sentences yield only a 4.4 percent gain over initial question-answering accuracy. Even though the correct answer was provided alongside the model's incorrect reasoning, the evaluated SLMs fail to understand what was missing in their reasoning and show minimal semantic difference between hints that lead to corrections and ones that do not. Furthermore, our experiments show that longer hints are positively correlated with incorrect final answers, suggesting that longer deliberation on problems can hinder the reasoning process, meaning that SLMs do not necessarily scale in performance with a larger compute budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。