arXiv:2606.19706cs.CVcs.CL2026-06被引 1

构建电影级长视频叙事事件结构数据集,助力理解跨时长情节关联。

NEST: Narrative Event Structures in Time for Long Video Understanding

论文配图:NEST: Narrative Event Structures in Time for Long Video Understanding
图 1 · 摘自论文原文
  • 基于视觉、对话与音频构建102个结构化叙事事件
  • 事件检测准确率低于8%,关系抽取达44.42%(微调后)
  • 适合研究长视频语义理解与情节推理的学者

视觉语言模型虽能处理长视频序列,却难以理解叙事结构。现有基准多聚焦于关键片段检索,而非低层动作如何构成事件、事件间如何随时间互动、情节如何推进。本文提出NEST数据集,包含1005部完整电影(平均98分钟),每部标注约102个融合视觉、对话和音频的多模态叙事事件,并通过时间顺序、层级结构与长程依赖关系建立事件关联。我们提供事件触发词检测(ETD)、事件定位(EL)、事件参数提取(EAE)和事件关系抽取(ERE)基线。该任务极具挑战:ETD低于8%,EL低于6%,EAE低于11%;而一旦事件给定,ERE在零样本下可达35.45% F1,微调后提升至44.42%。

原文摘要 · Abstract (English)

Recent progress in vision-language models has enabled the processing of increasingly long video sequences, but the ability to handle extended token streams does not translate to understanding of narrative structure in long videos. Existing long video benchmarks focus on needle-in-a-haystack retrieval rather than evaluating how low-level actions form events, how events interact across time, and how narratives progress, for example, whether a model can connect an early setback, such as a job loss to a later relationship breakup, despite long gaps, intervening scenes, or flashbacks that reframe what occurred. We introduce NEST (Narrative Event Structures in Time for Long Video Understanding), a dataset of 1005 full-length movies (avg. 98 minutes), each annotated with 102 multimodal narrative events grounded in visual content, dialogue, and audio. NEST captures multimodal narrative events with structured annotations grounded in visual content, dialogue, and audio, and links them through relations that reflect narrative structure, including temporal ordering, hierarchical composition, and long-range dependencies. We introduce baselines for event trigger detection (ETD), event localization (EL), event argument extraction (EAE), and event relation extraction (ERE). The benchmark is highly challenging for grounded event discovery, with ETD below 8%, EL under 6%, and EAE below 11%. In contrast, ERE is more tractable once events are given, reaching 35.45% F1 zero-shot and 44.42% F1 after fine-tuning.

长视频理解叙事结构事件抽取多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。