用事件链提升视频未来事件预测的逻辑推理能力
Video-CoE: Reinforcing Video Event Prediction via Chain of Events
- 构建事件链引导模型关注视觉内容与逻辑关联
- 在公开数据集上超越主流开源和商用模型
- 适合研究视频理解与因果推理的学者参考
尽管多模态大模型在多种视频任务中取得进展,视频事件预测(VEP)仍相对未被充分探索。VEP要求模型对视频进行细粒度的时间建模,并建立视频与未来事件间的逻辑关系,而现有多模态大模型仍难以胜任。本文首次对当前主流多模态大模型在VEP任务上的表现进行全面评估,揭示其预测不准的原因:缺乏对未来事件的逻辑推理能力,且未能充分利用视觉信息。为此,我们提出链式事件(Chain of Events, CoE)范式,通过构建时间事件链,隐式引导多模态大模型聚焦于视觉内容及视频与未来事件间的逻辑联系,并通过多种训练策略激励模型推理能力。在多个公开基准上的实验结果表明,该方法显著优于现有的领先开源与商业多模态大模型,建立了VEP任务的新基准。代码与模型将很快公开。
原文摘要 · Abstract (English)
Despite advances in the application of MLLMs for various video tasks, video event prediction (VEP) remains relatively underexplored. VEP requires the model to perform fine-grained temporal modeling of videos and establish logical relationships between videos and future events, which current MLLMs still struggle with. In this work, we first present a comprehensive evaluation of current leading MLLMs on the VEP task, revealing the reasons behind their inaccurate predictions, including lack of logical reasoning ability for future events prediction and insufficient utilization of visual information. To address these challenges, we propose \textbf{C}hain \textbf{o}f \textbf{E}vents (\textbf{CoE}) paradigm, which constructs temporal event chains to implicitly enforce MLLM focusing on the visual content and the logical connections between videos and future events, incentivizing model's reasoning capability with multiple training protocols. Experimental results on public benchmarks demonstrate that our method outperforms both leading open-source and commercial MLLMs, establishing a new state-of-the-art on the VEP task. Codes and models will be released soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。