RL提升视觉推理,关键在中后层计算的系统性优化
What does RL improve for Visual Reasoning? A Frankenstein-Style Analysis
- 通过因果探测定位改进部位,发现中后层是核心
- 参数对比显示中后层权重持续更新,且可迁移
- 适合研究多模态模型机制或评估训练策略的人
基于可验证奖励的强化学习(RL)已成为提升视觉语言模型视觉推理能力的标准后训练阶段,但其相比监督微调(IN)究竟提升了哪些具体能力仍不清晰。端到端基准性能提升掩盖了多个因素的混杂影响,难以归因。为此,本文提出一种类弗兰肯斯坦式分析框架,包含:(i) 基于因果探测的功能定位;(ii) 基于参数比较的更新表征;(iii) 基于模型合并的可迁移性测试。结果表明,RL主要引发中至后层推理过程的一致性变化,这些中后层的优化既可通过合并迁移,又在冻结时导致性能下降,证明其必要性。总体而言,RL在视觉推理中的可靠贡献并非对视觉感知的普遍增强,而是对中后层Transformer计算的系统性精炼,从而改善视觉到推理的对齐与推理表现,揭示了仅依赖基准评价难以真正理解多模态推理改进的局限。
原文摘要 · Abstract (English)
Reinforcement learning (RL) with verifiable rewards has become a standard post-training stage for boosting visual reasoning in vision-language models, yet it remains unclear what capabilities RL actually improves compared with supervised fine-tuning as cold-start initialization (IN). End-to-end benchmark gains conflate multiple factors, making it difficult to attribute improvements to specific skills. To bridge the gap, we propose a Frankenstein-style analysis framework including: (i) functional localization via causal probing; (ii) update characterization via parameter comparison; and (iii) transferability test via model merging. Instead, RL induces a consistent inference-time shift primarily in mid-to-late layers, and these mid-to-late refinements are both transferable (via merging) and necessary (via freezing) for RL gains. Overall, our results suggest that RL's reliable contribution in visual reasoning is not a uniform enhancement of visual perception, but a systematic refinement of mid-to-late transformer computation that improves vision-to-reasoning alignment and reasoning performance, highlighting the limitations of benchmark-only evaluation for understanding multimodal reasoning improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。