小模型通过自我反思学习,推理能力显著提升
ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection
- 设计迭代式反思训练流程,让小模型自动生成反思数据
- 在BIG-bench上将Llama-3准确率从52.4%提至71.2%
- 无需人工标注或大模型蒸馏,适合资源有限的研究者
我们提出一种新方法ReflectEvo,证明小语言模型(SLMs)可通过反思学习提升元认知能力。该过程通过迭代生成自我反思进行自训练,形成持续自我演进机制。基于此,我们构建了ReflectEvo-460k——一个大规模、全面的自生成反思数据集,涵盖多样化指令与多领域任务。利用该数据集,采用SFT和DPO训练,显著提升小模型推理能力:Llama-3准确率从52.4%提升至71.2%,Mistral从44.4%提升至71.1%。结果表明,ReflectEvo在不依赖大模型蒸馏或细粒度人工标注的前提下,可媲美甚至超越三个主流开源模型在BIG-bench上的推理表现。进一步分析显示,自生成反思具备高质量,有效助力错误定位与修正。本工作揭示了通过持续反思学习长期提升小模型推理性能的潜力。
原文摘要 · Abstract (English)
We present a novel pipeline, ReflectEvo, to demonstrate that small language models (SLMs) can enhance meta introspection through reflection learning. This process iteratively generates self-reflection for self-training, fostering a continuous and self-evolving process. Leveraging this pipeline, we construct ReflectEvo-460k, a large-scale, comprehensive, self-generated reflection dataset with broadened instructions and diverse multi-domain tasks. Building upon this dataset, we demonstrate the effectiveness of reflection learning to improve SLMs' reasoning abilities using SFT and DPO with remarkable performance, substantially boosting Llama-3 from 52.4% to 71.2% and Mistral from 44.4% to 71.1%. It validates that ReflectEvo can rival or even surpass the reasoning capability of the three prominent open-sourced models on BIG-bench without distillation from superior models or fine-grained human annotation. We further conduct a deeper analysis of the high quality of self-generated reflections and their impact on error localization and correction. Our work highlights the potential of continuously enhancing the reasoning performance of SLMs through iterative reflection learning in the long run.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。