让AI从零散视频中理解未完全定义的事件,提升多模态事件识别能力。
Grounding Partially-Defined Events in Multimodal Data
- 将部分定义事件建模为三阶段跨度检索任务,融合视频与文本信息。
- 构建包含14.5小时视频与22.8K实体标注的MultiVENT-G基准数据集。
- 基于大模型的方法在复杂事件理解上展现潜力,适合多模态系统研究者。
我们如何仅凭短视频片段就理解复杂的当前事件?虽然自然语言能直观表达不完整、部分可观测的事件,但视觉数据难以实现类似表示,因此在事件理解上带来独特挑战。随着具备视觉能力的AI代理日益普及,这些系统必须能从非结构化视频数据中建模事件。为此,我们提出一种多模态部分定义事件的形式化方法,并将事件提取建模为三阶段跨度检索任务。我们构建了名为MultiVENT-G的基准,包含14.5小时密集标注的实时事件视频和1,168篇文本文档,涵盖22.8K个以事件为中心的实体标注。我们提出一系列基于大模型的多模态事件分析方法,并在MultiVENT-G上进行评估。结果揭示了抽象事件理解的挑战,也展示了以事件为中心的视频-语言系统的发展前景。
原文摘要 · Abstract (English)
How are we able to learn about complex current events just from short snippets of video? While natural language enables straightforward ways to represent under-specified, partially observable events, visual data does not facilitate analogous methods and, consequently, introduces unique challenges in event understanding. With the growing prevalence of vision-capable AI agents, these systems must be able to model events from collections of unstructured video data. To tackle robust event modeling in multimodal settings, we introduce a multimodal formulation for partially-defined events and cast the extraction of these events as a three-stage span retrieval task. We propose a corresponding benchmark for this task, MultiVENT-G, that consists of 14.5 hours of densely annotated current event videos and 1,168 text documents, containing 22.8K labeled event-centric entities. We propose a collection of LLM-driven approaches to the task of multimodal event analysis, and evaluate them on MultiVENT-G. Results illustrate the challenges that abstract event understanding poses and demonstrates promise in event-centric video-language systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。