arXiv:2505.15447cs.CVcs.AI2025-05被引 10

用强化学习自动选视频关键帧,无需标注也能精准定位目标时刻。

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning

  • 通过迭代放大策略优化视频思维链中的帧选择过程
  • 在MLVU的针尖问答任务上提升近15%性能,优于现有方法
  • 适合需要高效理解长视频、追求零标注训练的应用场景

视频理解本质上是目标驱动的——人类会根据目的自然关注相关帧。尽管多模态大语言模型(MLLM)已实现灵活查询推理,但基于视频的思维链(Video CoT)缺乏直接训练信号来有效识别相关帧。现有方法多依赖启发式规则或伪标签监督,成本高且难以扩展。为此,我们提出首个基于规则的强化学习框架ViaRL,用于意图驱动的视频时序定位。采用迭代放大策略,在视频CoT系统中交替循环训练,使各组件不断优化。ViaRL利用下游模型的答案准确率作为奖励信号,通过试错训练帧选择器,无需昂贵标注,贴近人类学习过程。在VideoMME、LVBench和MLVU等多个基准上实验表明,ViaRL在时序定位性能与泛化能力上均显著领先,尤其在需从长视频中搜索特定目标的针尖问答(Needle QA)任务中提升近15%,验证了其有效性与可扩展性。

原文摘要 · Abstract (English)

Video understanding is inherently intention-driven-humans naturally focus on relevant frames based on their goals. Recent advancements in multimodal large language models (MLLMs) have enabled flexible query-driven reasoning; however, video-based frameworks like Video Chain-of-Thought lack direct training signals to effectively identify relevant frames. Current approaches often rely on heuristic methods or pseudo-label supervised annotations, which are both costly and limited in scalability across diverse scenarios. To overcome these challenges, we introduce ViaRL, the first framework to leverage rule-based reinforcement learning (RL) for optimizing frame selection in intention-driven video understanding. An iterated amplification strategy is adopted to perform alternating cyclic training in the video CoT system, where each component undergoes iterative cycles of refinement to improve its capabilities. ViaRL utilizes the answer accuracy of a downstream model as a reward signal to train a frame selector through trial-and-error, eliminating the need for expensive annotations while closely aligning with human-like learning processes. Comprehensive experiments across multiple benchmarks, including VideoMME, LVBench, and MLVU, demonstrate that ViaRL consistently delivers superior temporal grounding performance and robust generalization across diverse video understanding tasks, highlighting its effectiveness and scalability. Notably, ViaRL achieves a nearly 15\% improvement on Needle QA, a subset of MLVU, which is required to search a specific needle within a long video and regarded as one of the most suitable benchmarks for evaluating temporal grounding.

视频理解强化学习时序定位无标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。