arXiv:2604.01824cs.CV2026-04被引 1

通过结构化时空探索提升视频问答的强化学习性能

STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering

论文配图:STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
图 1 · 摘自论文原文
  • 构建视频时空变体并联合归一化文本与视觉输出
  • 在6个基准上超越强基线,显著提升推理稳定性
  • 适合需要鲁棒多视角推理的研究者和开发者

我们提出STRIVE(基于重要性感知变体探索的时空强化学习),一种用于视频问答的结构化强化学习框架。尽管基于分组的策略优化方法在大型多模态模型中表现良好,但当回答正确性相似时,常因奖励方差低导致优势估计弱或不稳定。STRIVE通过生成输入视频的多个时空变体,并对文本生成与视觉变体进行联合归一化,将群体比较从语言多样性扩展到结构化视觉扰动,从而丰富奖励信号,促进更稳定、更具信息量的策略更新。为确保探索语义合理,引入重要性感知采样机制,优先选择与问题相关的关键帧,同时保持时间覆盖。该设计促使模型在互补视觉视角间进行稳健推理,而非过拟合单一时空配置。在VideoMME、TempCompass、VideoMMMU、MMVU、VSI-Bench和PerceptionTest六个挑战性视频推理基准上的实验表明,STRIVE在多个大型多模态模型上持续优于强基线。结果凸显了结构化时空探索作为稳定多模态强化学习的原理性机制,能有效提升视频推理性能。

原文摘要 · Abstract (English)

We introduce STRIVE (SpatioTemporal Reinforcement with Importance-aware Variant Exploration), a structured reinforcement learning framework for video question answering. While group-based policy optimization methods have shown promise in large multimodal models, they often suffer from low reward variance when responses exhibit similar correctness, leading to weak or unstable advantage estimates. STRIVE addresses this limitation by constructing multiple spatiotemporal variants of each input video and performing joint normalization across both textual generations and visual variants. By expanding group comparisons beyond linguistic diversity to structured visual perturbations, STRIVE enriches reward signals and promotes more stable and informative policy updates. To ensure exploration remains semantically grounded, we introduce an importance-aware sampling mechanism that prioritizes frames most relevant to the input question while preserving temporal coverage. This design encourages robust reasoning across complementary visual perspectives rather than overfitting to a single spatiotemporal configuration. Experiments on six challenging video reasoning benchmarks including VideoMME, TempCompass, VideoMMMU, MMVU, VSI-Bench, and PerceptionTest demonstrate consistent improvements over strong reinforcement learning baselines across multiple large multimodal models. Our results highlight the role of structured spatiotemporal exploration as a principled mechanism for stabilizing multimodal reinforcement learning and improving video reasoning performance.

视频问答强化学习多模态时空探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。