让视觉语言模型真正基于图像证据推理,提升多模态强化学习的准确性。
SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning

- 通过局部图像替换构造干预样本,按实际效果而非操作类型分配监督信号。
- 在九个基准上显著提升3B/7B模型表现,视觉依赖任务提升8.79个百分点。
- 适用于多种强化学习框架,特别适合需要可靠视觉对齐的任务。
基于可验证奖励的强化学习(RLVR)推动多模态推理,但答案正确并不意味着模型真正基于视觉证据。现有视觉干预方法依据干预类型分配监督,但相同操作在不同样本中结果差异大。本文提出SIVA-RL框架,以样本级、结果导向的软监督替代操作条件监督。该方法通过像素对齐、距离受限的PatchSwap构建局部干预;冻结审计策略评估每对干净-干预图像,根据观察到的奖励下降生成软路由权重:大幅下降对对应敏感性对齐,小幅下降对对应清洁不变性对齐,模糊对则降权。该设计解耦了干预生成与监督分配,兼容GRPO与DAPO骨干。在涵盖数学、逻辑和视觉依赖任务的九个多模态推理基准上,SIVA-RL在所有设置下均优于匹配的强化学习基线,对3B与7B模型均有提升,在视觉依赖任务上取得8.79百分点增益,整体相对提升最高达14.9%。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards (RLVR) drives multimodal reasoning, but answer-level correctness does not guarantee that a vision-language model grounds its predictions in visual evidence. Existing visual-intervention methods contrast policy behavior on original and modified images, yet assign supervision by the type of intervention rather than its observed effect. This assumption fails: identical operators produce heterogeneous outcomes across samples. We propose SIVA-RL, a Sensitivity-Invariance Visual Alignment framework that replaces operator-conditioned regularization with sample-wise, outcome-conditioned supervision. SIVA-RL constructs localized interventions through token-aligned, distance-constrained within-image PatchSwap. A frozen audit policy then scores each clean-intervention pair, and the observed reward drop becomes soft routing weights. Large-drop pairs drive sensitivity alignment, low-drop pairs drive clean-anchored invariance alignment, and ambiguous pairs are down-weighted. This design decouples intervention construction from supervision assignment and is compatible with both GRPO and DAPO backbones. Across nine multimodal reasoning benchmarks spanning mathematical, logical, and vision-dependent tasks, SIVA-RL improves 3B and 7B models over matched RL baselines in every setting. It yields an 8.79 percentage-point gain on vision-dependent reasoning and up to 14.9% relative overall improvement across all four GRPO- and DAPO-based configurations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。