用两阶段方法让AI理解电影事件因果,更准地讲出故事来龙去脉。
Generating Event-oriented Attribution for Movies via Two-Stage Prefix-Enhanced Multimodal LLM
- 分两阶段:局部聚焦单片段,全局构建事件关联图
- 在两个真实数据集上超越现有最佳模型
- 适合做影视内容理解与智能叙事的开发者
社交媒体的兴起催生了对语义丰富服务的迫切需求,例如事件和剧情归因。然而,现有研究多集中于片段级事件理解,主要通过基础描述任务进行,缺乏对整部电影中事件成因的分析。这是一项重大挑战,因为即使先进的多模态大语言模型(MLLM)也受限于上下文长度,难以处理大量多模态信息。为此,我们提出一种两阶段前缀增强型多模态大语言模型(TSPE)方法,用于电影视频中的事件归因,即连接相关事件与其因果语义。在局部阶段,引入交互感知前缀,引导模型聚焦于单个片段内的相关多模态信息,简要总结单一事件;在全局阶段,通过推理知识图增强相关事件间的联系,并设计事件感知前缀,使模型关注相关事件而非所有先前片段,从而实现准确的事件归因。在两个真实世界数据集上的综合评估表明,该框架优于当前最先进方法。
原文摘要 · Abstract (English)
The prosperity of social media platforms has raised the urgent demand for semantic-rich services, e.g., event and storyline attribution. However, most existing research focuses on clip-level event understanding, primarily through basic captioning tasks, without analyzing the causes of events across an entire movie. This is a significant challenge, as even advanced multimodal large language models (MLLMs) struggle with extensive multimodal information due to limited context length. To address this issue, we propose a Two-Stage Prefix-Enhanced MLLM (TSPE) approach for event attribution, i.e., connecting associated events with their causal semantics, in movie videos. In the local stage, we introduce an interaction-aware prefix that guides the model to focus on the relevant multimodal information within a single clip, briefly summarizing the single event. Correspondingly, in the global stage, we strengthen the connections between associated events using an inferential knowledge graph, and design an event-aware prefix that directs the model to focus on associated events rather than all preceding clips, resulting in accurate event attribution. Comprehensive evaluations of two real-world datasets demonstrate that our framework outperforms state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。