调节奖励强度可控制视觉捷径的形成与消除,关键在干预时机。
When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR

- 用惩罚强度lambda作为控制变量,研究视觉捷径随训练时间的变化。
- 捷径在少数优化步内突然出现,且强度越高越能抑制捷径。
- 早期干预可阻止捷径形成,后期则效果显著下降。
基于可验证奖励的强化学习(RLVR)越来越多地应用于大规模视觉语言模型(LVLM),但仅依据结果优化可能导致模型放弃观看视频,转而依赖语言先验——我们称之为视觉捷径。尽管此类感知绕行现象已被记录,其形成机制、是否可逆及何时干预仍不明确。本文将接地惩罚强度lambda视为控制参数,刻画了视觉捷径在训练时间轴上的形成与逆转动态。在保留的、分布外的诊断数据集上发现:(i) 捷径依赖在极窄的优化步数窗口内突然出现,且对随机种子鲁棒;(ii) 呈单调剂量响应——提升lambda逐步抑制捷径,在中等强度下轨迹先形成后逆转,表现出获取与消除之间的类滞后不对称性;(iii) 存在关键干预窗口——在捷径出现前施加惩罚可有效阻止其形成,而在巩固后则效果大幅降低。这些结果表明,视觉捷径并非二元缺陷,而是可控、时变且不对称的过程,对多模态RLVR中何时及如何正则化具有直接启示。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards (RLVR) is increasingly applied to large vision-language models (LVLMs), yet outcome-only optimization can drive a model to stop attending to the video and instead exploit linguistic priors -- a failure we call a visual shortcut. While the existence of such perception bypass is by now documented, how it forms, whether it can be undone, and when intervention still helps remain open. We treat the strength of a grounding penalty, lambda, as a control knob and characterize the formation-reversal dynamics of visual shortcuts along the training time axis. On a held-out, out-of-distribution diagnostic set, we find: (i) a sharp onset -- shortcut reliance emerges abruptly over a narrow window of optimization steps and is robust across random seeds; (ii) a monotone dose-response -- increasing lambda progressively suppresses the shortcut, and at an intermediate dose the trajectory first forms and then reverses the shortcut, exposing a hysteresis-like asymmetry between acquiring and removing it; and (iii) a critical intervention window -- applying the penalty before onset arrests shortcut formation, whereas the same penalty applied after consolidation is markedly less effective. Together these results recast visual-shortcut collapse not as a binary defect but as a controllable, time-dependent, and asymmetric process, with direct implications for when and how strongly to regularize multimodal RLVR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。