评测视频叙事推理能力,挑战模型理解长时序因果链。
Narrative Aligned Long Form Video Question Answering
- 构建事件链记忆库,支持跨场景叙事信息整合
- 在远距离证据问题上,模型准确率仅42%,表现不佳
- 适合研究视频理解、叙事推理与多模态大模型的学者
多模态大模型在长视频推理任务上进展迅速,但现有基准多依赖局部线索,难以评估叙事推理能力——即追踪意图、连接遥远事件、重构整部影片的因果链条。我们提出NA-VQA,一个专为评估长视频深度时空与叙事推理设计的基准,包含88部完整电影和4400个开放式问答对,每题对应短、中、远三类证据跨度,用于检验长程依赖关系。通过要求生成跨多场景的答案,该基准测试模型是否能整合分散的叙事信息而非仅做浅层模式匹配。针对现有方法不足,我们提出Video-NaRA,一种以叙事为中心的框架,构建事件级链条并存储于结构化记忆中以供推理调用。大量实验表明,当前最先进多模态大模型在需远距离证据的问题上表现差(准确率仅42%),凸显显式叙事建模的必要性。Video-NaRA使长程推理性能提升最高达3个百分点,验证了其在处理复杂叙事结构上的有效性。我们将于论文发表后公开NA-VQA数据集。
原文摘要 · Abstract (English)
Recent progress in multimodal large language models (MLLMs) has led to a surge of benchmarks for long-video reasoning. However, most existing benchmarks rely on localized cues and fail to capture narrative reasoning, the ability to track intentions, connect distant events, and reconstruct causal chains across an entire movie. We introduce NA-VQA, a benchmark designed to evaluate deep temporal and narrative reasoning in long-form videos. NA-VQA contains 88 full-length movies and 4.4K open-ended question-answer pairs, each grounded in multiple evidence spans labeled as Short, Medium, or Far to assess long-range dependencies. By requiring generative, multi-scene answers, NA-VQA tests whether models can integrate dispersed narrative information rather than rely on shallow pattern matching. To address the limitations of existing approaches, we propose Video-NaRA, a narrative-centric framework that builds event-level chains and stores them in a structured memory for retrieval during reasoning. Extensive experiments show that state-of-the-art MLLMs perform poorly on questions requiring far-range evidence, highlighting the need for explicit narrative modeling. Video-NaRA improves long-range reasoning performance by up to 3 percent, demonstrating its effectiveness in handling complex narrative structures. We will release NA-VQA upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。