arXiv:2506.03340cs.CV2025-06NeurIPS被引 25

让大模型学会感知时间方向,提升视频理解能力。

Seeing the Arrow of Time in Large Multimodal Models

  • 用反向奖励强化学习,让模型区分正反视频帧的差异。
  • 在新基准上提升超20%准确率,标准视频问答也显著改善。
  • 适合关注视频时序理解、推理能力的研究者使用。

时间的不可逆性(时间之箭,AoT)是视频理解的基础,但现代大模型在处理语言查询时难以感知时间方向,阻碍了深层时序理解。本文首先分析现有评测基准与模型的不足,提出基于强化学习的ArrowRL训练策略,通过反向奖励促使模型对正序与倒序视频帧产生不同解读。为严格评估,我们构建了多维度的新基准AoTBench,用于探测复杂时序问题。实验表明,ArrowRL显著提升时序感知能力:在挑战性AoTBench上性能大幅提升,标准视频问答任务中最高准确率提升超过20%和10%。结果验证了该方法的有效性,并凸显了大模型中专门建模时间之箭的必要性。

原文摘要 · Abstract (English)

The Arrow of Time (AoT)-time's irreversible flow shaping physical events-is fundamental to video comprehension, yet remains a significant challenge for modern large multimodal models (LMMs). Current LMMs struggle to perceive and utilize temporal directionality in video when responding to language queries, obstructing deeper temporal understanding. We tackle this deficiency by first providing a critical analysis of existing benchmarks and models. We then introduce ArrowRL, a reinforcement learning (RL)-based training strategy with an innovative reverse reward that instills AoT awareness by encouraging divergent video interpretations between forward and reversed visual frames. For rigorous evaluation, we additionally develop AoTBench, a new multi-faceted benchmark probing temporally challenging questions. Experiments show ArrowRL greatly advances temporal perception: it not only achieves substantial improvements on our challenging AoTBench but also demonstrably boosts performance on standard video question answering (VQA) benchmarks (with peak accuracy gains reaching over 20% and 10% respectively). This validates ArrowRL's effectiveness and highlights the critical need for dedicated AoT understanding in LMMs.

时序理解大模型强化学习视频问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。