提出VISTA框架,精准预测长视频未来事件
Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining

- 分层次挖掘视频事件语义,从细节到叙事链逐步推进
- 在真实数据集上实现更准确的未来事件预测
- 适合需要长视频理解与推理的智能系统开发者
准确预测未来事件是内容理解与决策的核心。尽管已有研究集中于文本或短视频场景,但长视频因多模态上下文庞大、叙事复杂,仍鲜有探索。当前基于大语言模型和视觉语言模型的长视频语言模型虽在问答与摘要中表现良好,却难以泛化至事件预测,因其无法精确提取事件细节,也难进行细粒度分析。为此,本文提出VISTA——一种多层级事件语义挖掘框架。首先,采用以角色为中心的视觉提示,精准提取事件相关视觉细节,增强细节级语义;其次,通过知识增强的迭代检索策略,引导大模型逐步构建逻辑连贯的事件链,提升事件级叙事能力;最终,借鉴人类“先提出再检索”策略,生成多样化未来预测,并融合多层级线索,实现稳健且精准的预测。在真实数据集上的大量实验验证了VISTA在长视频事件预测中的有效性。
原文摘要 · Abstract (English)
Accurately predicting future events is fundamental to content understanding and decision-making across various domains. While prior research has primarily focused on text or short-video scenarios, long-video event prediction, characterized by vast multimodal context and more complex narratives, remains underexplored. Meanwhile, although recent Long-Video Language Models (LVLMs), built on Large Language Models (LLMs) and Vision-Language Models (VLMs), have shown promise in long-video question answering and summarization, they struggle to generalize to event prediction, as they can neither precisely extract event-related details nor perform fine-grained analysis of event development. To address this gap, we propose VISTA, a multi-level event semantics mining framework for long-video event prediction. Initially, VISTA applies a character-centric visual prompt to precisely extract event-related visual details, enhancing detail-level semantics; subsequently, it employs a knowledge-enhanced iterative retrieval strategy, guiding the LLM to progressively construct logically coherent event chains, thereby improving event-level narratives; ultimately, VISTA adopts a human-like propose-then-retrieve strategy to generate diverse future-oriented proposals and integrate multi-level clues, producing robust and accurate predictions. Extensive experiments on real-world datasets validate the effectiveness of VISTA for long-video event prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。