用长思维链监督微调,让小模型也能强推理
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
- 用长思维链数据做监督微调,提升小模型推理能力
- 微调后模型在推理任务上性能显著提升
- 适合想提升轻量多模态模型推理的开发者
尽管基于可验证奖励的强化学习已提升大模型的推理能力,但其在参数少于七亿的轻量级多模态语言模型(MLLM)中的有效性尚未充分探索。本文研究了长思维链(long CoT)数据对这类MLLM推理能力的增强作用。结果表明,使用长CoT数据进行监督微调(SFT)能显著提升轻量MLLM的推理表现。此外,在此SFT阶段之后,模型还能通过后续的强化学习阶段获得进一步性能提升。结论指出,以长CoT数据进行SFT是发展轻量级MLLM推理能力的关键前提。
原文摘要 · Abstract (English)
While Reinforcement Learning with Verifiable Rewards has enhanced the reasoning of large-scale language models (LLMs), its efficacy for lightweight multimodal language models (MLLMs) with fewer than seven billion parameters remains underexplored. This paper investigates the role of long Chain-of-Thought (long CoT) data in enhancing the reasoning abilities of such MLLMs. Our findings demonstrate that Supervised Fine-Tuning (SFT) with long CoT data significantly improves MLLM reasoning. Furthermore, we observe that after this initial SFT phase, MLLMs can achieve additional performance gains through a subsequent RL stage. We conclude that a SFT stage with long CoT data is a critical prerequisite for developing the reasoning capabilities of lightweight MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。