arXiv:2602.08503cs.CVcs.CL2026-02被引 3

通过重组合成纠错样本,提升视觉语言模型的自我修正能力。

Learning Self-Correction in Vision-Language Models via Rollout Augmentation

  • 用已有推理路径重组生成密集纠错数据,增强学习信号
  • 7个基准上达到开源模型最优,训练效率提升28%
  • 适合需要高精度推理的复杂视觉任务研究者

自我修正对视觉语言模型解决复杂推理问题至关重要。然而现有强化学习方法难以学习该能力,因有效纠错行为出现稀少,导致学习信号极度稀疏。为此,我们提出一种纠错专用回溯路径增强框架(Octopus),通过重组现有回溯路径合成密集的自我修正样本。该方法既提升了样本利用效率,又通过平衡监督稳定了强化学习优化过程。此外,引入响应掩码策略,将自我修正与直接推理解耦,避免信号冲突,实现两类行为的有效学习。基于此,我们构建了具备可控自我修正能力的Octopus-8B模型。在7个基准测试中,其性能超越现有开源视觉语言模型,优于最佳强化学习基线1.0分,且每步训练时间仅需其0.72倍。

原文摘要 · Abstract (English)

Self-correction is essential for solving complex reasoning problems in vision-language models (VLMs). However, existing reinforcement learning (RL) methods struggle to learn it, as effective self-correction behaviors emerge only rarely, making learning signals extremely sparse. To address this challenge, we propose correction-specific rollouts (Octopus), an RL rollout augmentation framework that synthesizes dense self-correction examples by recombining existing rollouts. This augmentation simultaneously improves sample efficiency due to rollout reuse and stabilizes RL optimization through balanced supervision. Furthermore, we introduce a response-masking strategy that decouples self-correction from direct reasoning, avoiding signal conflicts and enabling both behaviors to be learned effectively. Building on this, we introduce Octopus-8B, a reasoning VLM with controllable self-correction capability. Across 7 benchmarks, it achieves SoTA performance among open-source VLMs, outperforming the best RLVR baseline by 1.0 score while requiring only $0.72\times$ training time per step.

视觉语言模型强化学习自我修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。