解决多模态强化学习中思维与答案不一致的问题
CORA: Analyzing and bridging thinking-answer gap in Multimodal RLVR via Consistency-Oriented Reasoning Alignment

- 引入一致性奖励模型,对齐推理过程与最终答案语义
- 在多个基准上提升任务性能并减少推理偏差
- 适合关注多模态推理可靠性的研究者使用
基于可验证奖励的强化学习(RLVR)已成功激发大语言模型的推理能力,推动其向多模态场景拓展。现有方法主要关注推理轨迹的视觉覆盖和视觉幻觉缓解,却忽视了推理过程与最终答案之间的语义不一致问题。本文深入分析了在组相对策略优化(GRPO)训练及后置RLVR评估过程中收集的推演结果,发现该问题贯穿训练与推理全过程。为此,我们提出一致性导向的推理对齐方法(CORA),通过轻量级插件式一致性奖励模型引入思维-答案语义一致性,并结合混合奖励优势拆分(HRAS)稳定协调任务目标与一致性优化。在代表性多模态推理基准和主流大视觉语言模型上的大量实验表明,CORA在提升任务表现的同时有效缓解思维-答案不一致,生成更可信的推理轨迹。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards (RLVR) has successfully elicited the reasoning capabilities of large language models, motivating its extension to multimodal scenarios. Existing methods primarily focus on improving the visual coverage of reasoning traces and mitigating visual hallucinations, but underestimate the semantic inconsistency between the reasoning process and the final answer. In this paper, we delve into thinking-answer inconsistency in RLVR for large vision-language models (LVLMs), showing thorough analyses of rollouts collected throughout Group Relative Policy Optimization (GRPO) training process and post-RLVR evaluation outputs that this issue persists during training and remains present during inference. Motivated by the analysis, we propose Consistency-Oriented Reasoning Alignment (CORA), which introduces thinking-answer semantic consistency into RLVR through a lightweight plug-and-play consistency reward model, and further incorporates Hybrid Reward Advantage Splitting (HRAS) to stably coordinate task and consistency optimization. Extensive experiments across representative multimodal reasoning benchmarks and mainstream LVLMs show that CORA improves task performance while effectively mitigating thinking-answer inconsistency, leading to more faithful reasoning traces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。