构建首个面向长篇剧集多跳推理的基准,测试模型对跨集剧情的深层理解能力。
SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series

- 基于整季剧集设计跨集多跳推理任务,要求连接不同集之间的长期信息。
- 混合规划器在生成连贯推理链上表现最优,显著优于串行与并行策略。
- 适合研究视频叙事理解、多跳推理及智能体规划方法的学者和开发者。
我们提出SagaQA,一个针对完整电视剧长篇叙事进行多跳推理的视频基准。现有视频推理基准通常侧重相邻帧或片段的局部理解,而SagaQA填补了这一空白,要求对整部剧集的高阶多模态叙事具备全局理解。其核心特点是推理步骤的粒度:需在完全不同的剧集中建立信息关联,要求模型对事件与行为进行长期推理,实现对剧集叙述与进展的深层多模态理解。受近期智能体方法进展启发,我们进一步研究不同规划策略在复杂推理中的表现。将方法分为三类:并行、串行与混合规划器,并评估其生成连贯完整推理计划的能力。实验结果表明,混合规划器在生成高质量推理计划方面持续领先,展现出更强的复杂高阶叙事理解能力。
原文摘要 · Abstract (English)
We introduce SagaQA, a long-form video benchmark for multi-hop reasoning over full-length TV series. Existing video reasoning benchmarks often emphasize local understanding of adjacent frames or clips. SagaQA addresses this gap by requiring high-level comprehension of extended multimodal narratives in entire TV shows. A distinguishing feature of SagaQA is the granularity of its reasoning steps. Our dataset necessitates long-range reasoning hops to connect information across completely different episodes. This requires models to reason over entire events and actions, demanding a deep understanding of the show's narration and progression at a multimodal level. Motivated by recent progress in agentic methods, we further study how different planning strategies handle such complex reasoning. We categorize these approaches into three classes-Parallel, Sequential, and Hybrid planners-and evaluate their ability to generate coherent and complete reasoning plans. Our results on SagaQA suggest that hybrid planners consistently produce higher-quality plans and exhibit stronger capabilities for complex, high-level narrative understanding in TV shows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。