用推理模型的答案提升普通模型的问答能力
Leveraging Reasoning Model Answers to Enhance Non-Reasoning Model Capability
- 用强推理模型生成答案,指导弱模型学习
- 在多个基准上实现稳定性能提升
- 适合资源有限但需高质量问答的场景
近期大语言模型(如 DeepSeek-R1 和 OpenAI-o1)通过测试时扩展(test-time scaling)展现出显著性能提升,其通过系统性“思考”步骤提高答案质量。本文提出利用这些高精度推理模型输出的答案,来增强计算成本更低的非推理模型能力。我们探索并比较了多种利用推理模型输出训练和优化非推理模型的方法。在主流基准上的简单监督微调(SFT)实验表明,该方法在多个任务中均实现一致性能提升,验证了此策略在直接提升模型问答能力方面的潜力。
原文摘要 · Abstract (English)
Recent advancements in large language models (LLMs), such as DeepSeek-R1 and OpenAI-o1, have demonstrated the significant effectiveness of test-time scaling, achieving substantial performance gains across various benchmarks. These advanced models utilize deliberate "thinking" steps to systematically enhance answer quality. In this paper, we propose leveraging these high-quality outputs generated by reasoning-intensive models to improve less computationally demanding, non-reasoning models. We explore and compare methodologies for utilizing the answers produced by reasoning models to train and improve non-reasoning models. Through straightforward Supervised Fine-Tuning (SFT) experiments on established benchmarks, we demonstrate consistent improvements across various benchmarks, underscoring the potential of this approach for advancing the ability of models to answer questions directly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。