构建首个剧情驱动剧集理解基准,推动多模态模型读懂连续叙事。
SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
- 设计长跨度剧情标注方法,将人工标注转为28类任务
- 在105部剧集中验证,现有模型仍难理解复杂剧情
- 提出PC-DCoT框架,提升角色关系与情节结构分析能力
随着多模态大模型的快速发展,大量视频理解评测基准被建立,但多数聚焦单个视频,主要评估人物动作、物体状态等视觉元素。现实中,当代视频常以连续剧集形式呈现复杂叙事。为此,我们提出SeriesBench,一个包含105部精心挑选的剧情驱动剧集的基准,涵盖28项需深度叙事理解的任务。首先,选取跨类型多样化的戏剧剧集;其次,引入新型长跨度叙事标注方法,并结合全信息转换策略,将人工标注转化为多种任务格式。为进一步提升模型对剧集情节结构与角色关系的细粒度分析能力,我们提出新型叙事推理框架PC-DCoT。SeriesBench上的广泛实验表明,现有多模态大模型在理解剧情驱动剧集方面仍面临显著挑战,而PC-DCoT可有效提升其性能。整体上,SeriesBench与PC-DCoT凸显了增强模型对剧情驱动剧集理解能力的紧迫性,为未来多模态大模型的发展提供指引。数据集已公开于https://github.com/zackhxn/SeriesBench-CVPR2025。
原文摘要 · Abstract (English)
With the rapid development of Multi-modal Large Language Models (MLLMs), an increasing number of benchmarks have been established to evaluate the video understanding capabilities of these models. However, these benchmarks focus on standalone videos and mainly assess "visual elements" like human actions and object states. In reality, contemporary videos often encompass complex and continuous narratives, typically presented as a series. To address this challenge, we propose SeriesBench, a benchmark consisting of 105 carefully curated narrative-driven series, covering 28 specialized tasks that require deep narrative understanding. Specifically, we first select a diverse set of drama series spanning various genres. Then, we introduce a novel long-span narrative annotation method, combined with a full-information transformation approach to convert manual annotations into diverse task formats. To further enhance model capacity for detailed analysis of plot structures and character relationships within series, we propose a novel narrative reasoning framework, PC-DCoT. Extensive results on SeriesBench indicate that existing MLLMs still face significant challenges in understanding narrative-driven series, while PC-DCoT enables these MLLMs to achieve performance improvements. Overall, our SeriesBench and PC-DCoT highlight the critical necessity of advancing model capabilities to understand narrative-driven series, guiding the future development of MLLMs. SeriesBench is publicly available at https://github.com/zackhxn/SeriesBench-CVPR2025.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。