探究长思维链与强化学习在多模态推理中的协同困境
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
- 用长思维链微调提升难题解答能力,但导致回答冗长
- 强化学习增强泛化性与简洁性,对难题提升有限
- 混合训练反而引发性能折衷,难实现协同增益
大型视觉语言模型(VLMs)越来越多地采用后训练技术,如长链式思维(CoT)监督微调(SFT)和强化学习(RL),以激发复杂推理能力。尽管这些方法在纯语言模型中表现出协同效应,但在多模态模型中的联合效果仍不明确。本文系统研究了长CoT SFT与RL在多个跨模态推理基准上的作用与相互关系。结果发现,SFT通过深入、结构化的推理提升难题表现,但造成回答冗长,并损害简单问题的性能;而RL促进泛化性和简洁性,在所有难度级别上均带来稳定提升,但在最难问题上的增益不如SFT显著。令人意外的是,通过两阶段、交替或渐进式训练、数据混合及模型融合等策略结合二者,均未能产生叠加优势,反而在准确率、推理风格和响应长度间引入权衡。这一‘协同困境’凸显出亟需更无缝、自适应的方法来释放联合后训练技术在推理型VLM中的全部潜力。
原文摘要 · Abstract (English)
Large vision-language models (VLMs) increasingly adopt post-training techniques such as long chain-of-thought (CoT) supervised fine-tuning (SFT) and reinforcement learning (RL) to elicit sophisticated reasoning. While these methods exhibit synergy in language-only models, their joint effectiveness in VLMs remains uncertain. We present a systematic investigation into the distinct roles and interplay of long-CoT SFT and RL across multiple multimodal reasoning benchmarks. We find that SFT improves performance on difficult questions by in-depth, structured reasoning, but introduces verbosity and degrades performance on simpler ones. In contrast, RL promotes generalization and brevity, yielding consistent improvements across all difficulty levels, though the improvements on the hardest questions are less prominent compared to SFT. Surprisingly, combining them through two-staged, interleaved, or progressive training strategies, as well as data mixing and model merging, all fails to produce additive benefits, instead leading to trade-offs in accuracy, reasoning style, and response length. This ``synergy dilemma'' highlights the need for more seamless and adaptive approaches to unlock the full potential of combined post-training techniques for reasoning VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。