构建了迄今最大的深度视频理解数据集,提升复杂剧情理解能力。
StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset

- 设计多智能体协作框架自动生成高质量问答对
- 建成363万条问答、393小时跨类型视频数据集
- 提出剧情树结构模型,支持长程剧情推理
视频问答(VideoQA)旨在回答给定视频相关的问题。现有方法在事实型问答上表现良好,但在需要理解复杂剧情的深度视频理解(DVU)任务上表现不佳,原因在于视频内容长时序、问题类型多样、情节细节精细,且人工构建的DVU数据集规模与多样性受限。为此,我们此前提出StoryMind,可自动生成电视剧的细粒度话题问答对,但在处理更长更复杂的电影时性能显著下降。本文进一步设计StoryMindv2,一种增强的多智能体协作框架,可生成电视系列剧和电影的高质量DVU数据集。通过引入监督引导生成机制和优化的多评审投票策略,构建了目前最大的DVU数据集StoryVideoQA,包含超过36.3万条问答,覆盖393.2小时多样剧情视频,涵盖电视剧(平均1,635秒)和电影(平均7,878秒)。对20种先进VideoQA模型在该大规模基准上的评估显示,它们难以维持长程角色关联或构建连贯的剧情理解。为弥补这一差距,我们提出PlotTree,一种新型视频理解代理,将长时序视频内容重组为层次化剧情结构,实现对StoryVideoQA的有效剧情推理。
原文摘要 · Abstract (English)
Video question answering (VideoQA) aims to answer questions about given videos. While existing approaches excel on factoid VideoQA, they struggle with deep video understanding (DVU), which requires the comprehension of complex storylines. This challenge arises from the inherent long-range video content, multi-faceted question types, and instance-level story elements, all of which constrain the scale and diversity of manually constructed DVU datasets. These difficulties constrain the scale and diversity of manually-constructed DVU dataset. To address these, we previously introduced StoryMind to automatically construct DVU datasets with balanced fine-grained topics. Though it can generate high-quality question-answer pairs (QAs) for TV series, it suffers significant performance degradation when handling longer and more complex movies. In this paper, we further design StoryMindv2, an enhanced multi-agent collaboration framework to generate high-quality DVU datasets for both TV series and movies. By integrating a novel supervisor-guided generation mechanism and a refined multi-reviewer voting strategy, the framework is utilized to construct StoryVideoQA, the largest DVU dataset to date, featuring over 363K QAs on 393.2 hours diverse story videos including TV series (avg. 1,635 seconds) and movies (avg. 7,878 seconds). Comprehensive evaluations of 20 state-of-the-art VideoQA methods on this large-scale benchmark reveal that they cannot fully maintain long-range character associations or construct a coherent understanding of complex storylines. To bridge this gap, we propose PlotTree, a novel video understanding agent, re-organizing long-range video content into a hierarchical plot structure, enabling efficient storyline reasoning on StoryVideoQA. Project page: https://github.com/nercms-mmap/StoryVideoQA/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。