为视觉语言模型设计细粒度推理链评估,提升推理质量与训练效果。
Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- 构建逐步推理数据与过程奖励模型,实现中间推理步骤的精准评估。
- 在多个复杂视觉语言基准上取得显著提升,验证方法有效性。
- 适合研究多模态推理、强化学习与可解释性模型的学者参考。
思维链推理在大语言模型中表现卓越,但其在视觉语言模型中的应用仍面临挑战,现有方法多采用粗粒度推理链,难以实现精细结构化推理,且难以评估中间步骤的质量。本文深入探索视觉语言模型的逐步推理机制,提出一种简单、高效且完全透明的框架,包含步骤级推理数据、过程奖励模型(PRM)和强化学习训练流程。该框架可准确评估每一步推理质量,推动有效强化学习与推理时缩放。实验表明,所提模型在多个挑战性视觉语言基准上建立强基线,持续提升性能。此外,我们进行了详尽的实证分析与消融研究,揭示各组件影响及推理时缩放的若干有趣特性。我们认为本工作为视觉语言模型提供了基础范式,并为更复杂的多模态推理提供洞见。数据集、PRM与代码将公开于 https://github.com/baaivision/CoS。
原文摘要 · Abstract (English)
Chain of thought reasoning has demonstrated remarkable success in large language models, yet its adaptation to vision-language reasoning remains an open challenge with unclear best practices. Existing attempts typically employ reasoning chains at a coarse-grained level, which struggles to perform fine-grained structured reasoning and, more importantly, are difficult to evaluate the reward and quality of intermediate reasoning. In this work, we delve into chain of step reasoning for vision-language models, enabling assessing reasoning step quality accurately and leading to effective reinforcement learning and inference-time scaling with fine-grained rewards. We present a simple, effective, and fully transparent framework, including the step-level reasoning data, process reward model (PRM), and reinforcement learning training. With the proposed approaches, our models set strong baselines with consistent improvements on challenging vision-language benchmarks. More importantly, we conduct a thorough empirical analysis and ablation study, unveiling the impact of each component and several intriguing properties of inference-time scaling. We believe this paper serves as a baseline for vision-language models and offers insights into more complex multimodal reasoning. Our dataset, PRM, and code will be available at https://github.com/baaivision/CoS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。