arXiv:2412.17022cs.CV2024-12中稿 · AAAI被引 3

构建首个基于细粒度话题分类的长视频理解数据集,助力模型理解剧情主线。

FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story Videos

  • 用大模型多智能体框架自动生成剧情问答数据
  • 含44.6万题,覆盖14类故事主题,每集平均时长1358秒
  • 适合评估模型对复杂剧情线索的理解能力

视频问答(VideoQA)旨在根据视频内容回答自然语言问题。尽管现有模型在事实型视频问答任务中表现良好,但在深度视频理解(DVU)任务上仍面临挑战,尤其针对包含复杂情节的剧情类视频。与事实型视频相比,剧情视频的核心特征是故事情节,由角色、动作、地点等核心话题的复杂互动和长期演变构成。理解这些话题需要模型具备深度视频理解能力。然而,现有DVU数据集通常未按故事话题组织问题,难以全面评估模型对复杂情节的理解能力。此外,由于人工标注成本高,现有数据集在问题数量和视频长度上均受限。本文提出基于大语言模型的多智能体协作框架StoryMind,自动构建新的大规模DVU数据集FriendsQA。该数据集源自著名情景剧《老友记》,平均每集时长1358秒,共包含44.6万道问题,均匀分布于14个细粒度话题类别。我们进一步在10个主流VideoQA模型上进行了全面实验,验证了该数据集的有效性。

原文摘要 · Abstract (English)

Video question answering (VideoQA) aims to answer natural language questions according to the given videos. Although existing models perform well in the factoid VideoQA task, they still face challenges in deep video understanding (DVU) task, which focuses on story videos. Compared to factoid videos, the most significant feature of story videos is storylines, which are composed of complex interactions and long-range evolvement of core story topics including characters, actions and locations. Understanding these topics requires models to possess DVU capability. However, existing DVU datasets rarely organize questions according to these story topics, making them difficult to comprehensively assess VideoQA models' DVU capability of complex storylines. Additionally, the question quantity and video length of these dataset are limited by high labor costs of handcrafted dataset building method. In this paper, we devise a large language model based multi-agent collaboration framework, StoryMind, to automatically generate a new large-scale DVU dataset. The dataset, FriendsQA, derived from the renowned sitcom Friends with an average episode length of 1,358 seconds, contains 44.6K questions evenly distributed across 14 fine-grained topics. Finally, We conduct comprehensive experiments on 10 state-of-the-art VideoQA models using the FriendsQA dataset.

视频理解剧情分析大模型生成数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。