通过修复错误推理轨迹,自动生成多模态过程监督,提升视觉推理对齐效果。
Difference Feedback: Generating Multimodal Process-Level Supervision for VLM Reinforcement Learning
- 用错误修复法自动生成需纠正的步骤标记
- 在多模态推理任务上平均提升3%准确率
- 无需人工标注,可直接接入现有强化学习框架
视觉-语言模型(VLM)正通过类似组相对策略优化(GRPO)的方法进行对齐。然而,仅依赖最终结果奖励会导致多步推理中的信用分配稀疏,削弱视觉证据与中间步骤的关联,常引发优化不稳定和视觉幻觉。我们提出差分反馈(Difference Feedback),通过修复错误推理轨迹,自动构建逐标记/步骤级的监督掩码,明确标出需修正的关键位置。无需昂贵的大规模逐步人工标注,该方法实现了过程级视觉对齐,且可无缝集成到现有GRPO类框架中。在包含MMMStar和MathVista在内的多模态推理基准上,相同计算预算下平均提升3%。本方法为精确的视觉-推理过程对齐提供了高效、低成本的解决方案。
原文摘要 · Abstract (English)
Vision--language models (VLMs) are increasingly aligned via Group Relative Policy Optimization (GRPO)-style training. However, relying solely on terminal outcome rewards yields sparse credit assignment in multi-step reasoning, weakening the linkage between visual evidence and intermediate steps and often causing unstable optimization and visual hallucinations. We propose Differential Feedback, which automatically constructs token/step-level supervision masks by repairing erroneous reasoning trajectories, explicitly marking the key positions that require correction. Without costly large-scale step-by-step human annotations, our method enables process-level visual alignment and can be seamlessly integrated into existing GRPO-like frameworks. Experiments on multimodal reasoning benchmarks including MMMStar and MathVista show an average 3% improvement under matched compute budgets. Our approach offers an effective, low-cost solution for accurate vision--reasoning process alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。