arXiv:2509.03321cs.CV2025-09被引 1

用长思维链监督微调,让小模型也能强推理

Empowering Lightweight MLLMs with Reasoning via Long CoT SFT

  • 用长思维链数据做监督微调,提升小模型推理能力
  • 微调后模型在推理任务上性能显著提升
  • 适合想提升轻量多模态模型推理的开发者

尽管基于可验证奖励的强化学习已提升大模型的推理能力,但其在参数少于七亿的轻量级多模态语言模型(MLLM)中的有效性尚未充分探索。本文研究了长思维链(long CoT)数据对这类MLLM推理能力的增强作用。结果表明,使用长CoT数据进行监督微调(SFT)能显著提升轻量MLLM的推理表现。此外,在此SFT阶段之后,模型还能通过后续的强化学习阶段获得进一步性能提升。结论指出,以长CoT数据进行SFT是发展轻量级MLLM推理能力的关键前提。

原文摘要 · Abstract (English)

While Reinforcement Learning with Verifiable Rewards has enhanced the reasoning of large-scale language models (LLMs), its efficacy for lightweight multimodal language models (MLLMs) with fewer than seven billion parameters remains underexplored. This paper investigates the role of long Chain-of-Thought (long CoT) data in enhancing the reasoning abilities of such MLLMs. Our findings demonstrate that Supervised Fine-Tuning (SFT) with long CoT data significantly improves MLLM reasoning. Furthermore, we observe that after this initial SFT phase, MLLMs can achieve additional performance gains through a subsequent RL stage. We conclude that a SFT stage with long CoT data is a critical prerequisite for developing the reasoning capabilities of lightweight MLLMs.

多模态轻量模型思维链推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。