arXiv:2603.10652cs.CVcs.AI2026-03中稿 · ECCV被引 2

让视频推理模型在真实环境扰动下仍能稳定工作。

Are Video Reasoning Models Ready to Go Outside?

  • 通过自反思评估动态识别样本难度,实现鲁棒性感知的训练优化。
  • 在真实扰动下模型准确率与推理能力下降最多达35%和28%,ROVA提升超24%和9%。
  • 新基准PVRBench模拟现实干扰,适合评估视频模型真实部署性能。

在真实部署中,视觉语言模型常面临天气、遮挡和摄像机运动等干扰,导致理解与推理能力显著下降,暴露出干净可控测试与真实鲁棒性之间的差距。为此,我们提出ROVA训练框架,通过建模时空扰动下的鲁棒性感知一致性奖励,提升模型抗干扰能力。ROVA采用难度感知的在线训练策略,基于模型动态能力持续重估样本难度,实现自适应训练。我们还构建了新基准PVRBench,将真实世界扰动注入具身视频数据集,以评估在现实干扰下的准确率与推理质量。在PVRBench、UrbanVideo和VisBench上,开源与专有模型在真实扰动下准确率与推理能力最高分别下降35%和28%。ROVA有效缓解性能衰减,相较基线模型(QWen2.5/3-VL, InternVL2.5, Embodied-R)相对准确率提升至少24%,推理能力提升超9%。该增益亦可迁移至标准清洁基准,实现一致改进。

原文摘要 · Abstract (English)

In real-world deployment, vision-language models often encounter disturbances such as weather, occlusion, and camera motion. Under such conditions, their understanding and reasoning degrade substantially, revealing a gap between clean, controlled (i.e., unperturbed) evaluation settings and real-world robustness. To address this limitation, we propose ROVA, a novel training framework that improves robustness by modeling a robustness-aware consistency reward under spatio-temporal corruptions. ROVA introduces a difficulty-aware online training strategy that prioritizes informative samples based on the model's evolving capability. Specifically, it continuously re-estimates sample difficulty via self-reflective evaluation, enabling adaptive training with a robustness-aware consistency reward. We also introduce PVRBench, a new benchmark that injects real-world perturbations into embodied video datasets to assess both accuracy and reasoning quality under realistic disturbances. We evaluate ROVA and baselines on PVRBench, UrbanVideo, and VisBench, where open-source and proprietary models suffer up to 35% and 28% drops in accuracy and reasoning under realistic perturbations. ROVA effectively mitigates performance degradation, boosting relative accuracy by at least 24% and reasoning by over 9% compared with baseline models (QWen2.5/3-VL, InternVL2.5, Embodied-R). These gains transfer to clean standard benchmarks, yielding consistent improvements.

视频推理鲁棒性训练框架真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。