对比强化学习与监督微调,发现视觉语言模型在跨模态推理中存在组合能力短板。
Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model
- 设计多模态组合任务,测试模型跨任务、跨模态的技能整合能力。
- 强化学习训练模型在组合泛化上显著优于监督微调,但整体仍表现不足。
- 先描述图像再推理+逐步视觉到文本对齐,可显著提升组合推理效果。
尽管大语言模型通过可验证奖励的强化学习展现出强大推理能力,但大型视觉语言模型(VLMs)是否可通过类似后训练策略继承此类能力仍待探索。本文系统性地开展组合性探测研究,评估当前采用强化学习或其他后训练策略的VLMs在分布外条件下跨模态或跨任务的组合能力。设计了一系列诊断任务,让模型先学习单模态任务或孤立推理技能,再评估其在需技能融合的多模态组合任务上的表现。通过对比监督微调(SFT)与强化学习(RL)训练模型,发现:(1) 强化学习训练模型在组合泛化上持续优于监督微调,展现更好技能整合能力;(2) 尽管单任务表现优异,现有VLMs在跨模态、跨任务场景下仍严重缺乏组合泛化能力,暴露当前训练策略的重大缺陷;(3) 要求模型在推理前显式描述视觉内容(如‘先生成图像描述’),并奖励渐进式的视觉到文本对齐,能带来显著性能提升。结果表明,视觉到文本对齐与准确视觉定位是提升VLM组合性的两个关键要素。研究揭示了基于强化学习的推理型VLM训练的局限,并为构建跨模态、跨任务组合推理模型提供可行路径。
原文摘要 · Abstract (English)
While large language models (LLMs) demonstrate strong reasoning capabilities utilizing reinforcement learning (RL) with verifiable reward, whether large vision-language models (VLMs) can directly inherit such capabilities through similar post-training strategies remains underexplored. In this work, we conduct a systematic compositional probing study to evaluate whether current VLMs trained with RL or other post-training strategies can compose capabilities across modalities or tasks under out-of-distribution conditions. We design a suite of diagnostic tasks that train models on unimodal tasks or isolated reasoning skills, and evaluate them on multimodal, compositional variants requiring skill integration. Through comparisons between supervised fine-tuning (SFT) and RL-trained models, we identify three key findings: (1) RL-trained models consistently outperform SFT on compositional generalization, demonstrating better integration of learned skills; (2) although VLMs achieve strong performance on individual tasks, they struggle to generalize compositionally under cross-modal and cross-task scenario, revealing a significant gap in current training strategies; (3) enforcing models to explicitly describe visual content before reasoning (e.g., caption-before-thinking), along with rewarding progressive vision-to-text grounding, yields notable gains. It highlights two essential ingredients for improving compositionality in VLMs: visual-to-text alignment and accurate visual grounding. Our findings shed light on the current limitations of RL-based reasoning VLM training and provide actionable insights toward building models that reason compositionally across modalities and tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。