arXiv:2509.23433cs.CVcs.CL2025-09

让视频大模型学会识别意外事件,提升对关键片段的捕捉能力。

SPIKE-RL: Video-LLMs meet Bayesian Surprise

  • 用贝叶斯意外度量化视觉信息与已有认知的冲突,定位视频中的关键瞬间。
  • 在FunQA和Oops!数据集上与人类判断高度相关,显著优于均匀采样。
  • 适用于需要关注重要事件的场景,如视频理解、自动剪辑和异常检测。

现实世界视频常由常规活动穿插着令人印象深刻的意外事件组成。然而,多数视频大模型采用均匀采样帧,可能遗漏定义叙事的关键时刻。我们提出SPIKE,一种推理时框架,通过量化贝叶斯意外度(即新视觉证据引发的认知更新)来识别与先前信念冲突的时刻。SPIKE能有效定位视频中的意外事件,在FunQA(正面意外)和Oops!(负面意外)基准上与人类判断高度相关。由于零样本视频大模型的信念常不准确,我们进一步开发SPIKE-RL,利用GRPO优化信念假设,以视频字幕作为奖励信号。SPIKE与SPIKE-RL引导无查询依赖的惊喜加权帧采样,将更多帧分配给有趣片段。该策略在五个下游任务上持续优于均匀采样,显著提升性能。本工作使视频大模型具备跟踪信念与记录意外的能力,为可动态修正理解的鲁棒模型开辟道路。

原文摘要 · Abstract (English)

Real-world videos often show routine activities punctuated by memorable, surprising events. However, most Video-LLMs process videos by sampling frames uniformly, likely missing critical moments that define a video's narrative. We introduce SPIKE, an inference-time framework that quantifies Bayesian Surprise as the belief update triggered by new visual evidence in the video stream, identifying moments where new visual evidence conflicts with prior beliefs. SPIKE effectively localizes surprise in videos, strongly correlated with humans on positive (FunQA) and negative (Oops!) surprise benchmarks. Since the beliefs of zero-shot Video-LLMs are often suboptimal, we develop SPIKE-RL, which leverages GRPO to optimize belief hypotheses based on a reward signal from the video caption. SPIKE and SPIKE-RL guide query-agnostic surprise-weighted frame sampling, which allocates more frames to interesting moments in the video. With this strategy, we achieve consistent performance gains on five downstream benchmarks over uniform sampling. By enabling Video-LLMs to track beliefs and register surprise, our work paves the way for more robust models that can revise their understanding in response to new information.

视频理解贝叶斯推理意外检测帧采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。