arXiv:2510.02240cs.CVcs.AI2025-10被引 26

用多阶段强化学习解决视觉推理中的奖励稀疏问题。

RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning

  • 分阶段训练,从感知到推理逐步提升模型能力。
  • 引入难度感知奖励机制,实现更丰富的监督信号。
  • 在6个基准上平均提升3.47%,适用于复杂视觉推理任务。

细粒度视觉推理仍是多模态大语言模型(MLLMs)的核心挑战。近期提出的ReasonMap揭示了即使先进MLLMs在交通图等结构化、信息密集场景下的空间推理能力仍不足,该任务具有重要的实际与科学意义。然而,标准强化学习在此类任务中受制于奖励稀疏和优化不稳。为此,我们构建了ReasonMap-Plus数据集,通过视觉问答(VQA)任务引入密集奖励信号,支持细粒度视觉理解技能的冷启动训练。进一步提出RewardMap多阶段强化学习框架,包含两项关键设计:一是难度感知奖励机制,融合细节奖励以缓解奖励稀疏并提供更丰富监督;二是多阶段训练策略,从简单感知逐步过渡到复杂推理,相比传统监督微调(SFT)提供更有效的冷启动方案。在ReasonMap与ReasonMap-Plus上的实验表明,每个组件均带来稳定性能提升,组合效果最佳。此外,使用RewardMap训练的模型在涵盖空间推理、细粒度视觉推理及跨地图通用任务的6个基准上,平均提升3.47%,验证了其在增强视觉理解与推理能力方面的有效性。

原文摘要 · Abstract (English)

Fine-grained visual reasoning remains a core challenge for multimodal large language models (MLLMs). The recently introduced ReasonMap highlights this gap by showing that even advanced MLLMs struggle with spatial reasoning in structured and information-rich settings such as transit maps, a task of clear practical and scientific importance. However, standard reinforcement learning (RL) on such tasks is impeded by sparse rewards and unstable optimization. To address this, we first construct ReasonMap-Plus, an extended dataset that introduces dense reward signals through Visual Question Answering (VQA) tasks, enabling effective cold-start training of fine-grained visual understanding skills. Next, we propose RewardMap, a multi-stage RL framework designed to improve both visual understanding and reasoning capabilities of MLLMs. RewardMap incorporates two key designs. First, we introduce a difficulty-aware reward design that incorporates detail rewards, directly tackling the sparse rewards while providing richer supervision. Second, we propose a multi-stage RL scheme that bootstraps training from simple perception to complex reasoning tasks, offering a more effective cold-start strategy than conventional Supervised Fine-Tuning (SFT). Experiments on ReasonMap and ReasonMap-Plus demonstrate that each component of RewardMap contributes to consistent performance gains, while their combination yields the best results. Moreover, models trained with RewardMap achieve an average improvement of 3.47% across 6 benchmarks spanning spatial reasoning, fine-grained visual reasoning, and general tasks beyond transit maps, underscoring enhanced visual understanding and reasoning capabilities.

视觉推理强化学习多阶段训练奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。