arXiv:2603.27259cs.CV2026-03中稿 · CVPR

新基准发现视频模型在长时场景理解中严重遗忘,需增强上下文记忆。

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark

论文配图:Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark
图 1 · 摘自论文原文
  • 以场景为单位构建长视频评测基准,贴近人类感知。
  • 模型在场景级问题上准确率大幅下降,暴露长期记忆缺陷。
  • 提出动态场景记忆机制,提升性能2.50%,适合研究长视频模型者参考。

长视频理解(LVU)仍是多模态学习的核心挑战。尽管近期视觉-语言模型(VLMs)取得显著进展,现有基准主要集中于细粒度感知或粗粒度摘要,难以揭示长时程上下文的理解能力。本文定义场景为视觉与语义上下文保持一致的视频连贯片段,契合人类感知。由此提出核心问题:当前VLM能否有效推理长时场景级上下文?为此,我们引入新基准SceneBench,专攻场景级挑战。评估显示,当VLM回答场景级问题时,准确率急剧下降,表明存在显著的长程上下文遗忘。为进一步验证,我们提出场景检索增强生成(Scene-RAG),通过跨场景检索与整合构建动态场景记忆。该方法使模型性能提升+2.50%,证实当前模型仍难以维持长时上下文。我们期望SceneBench能推动面向更鲁棒、类人视频理解的VLM研究。

原文摘要 · Abstract (English)

Long video understanding (LVU) remains a core challenge in multimodal learning. Although recent vision-language models (VLMs) have made notable progress, existing benchmarks mainly focus on either fine-grained perception or coarse summarization, offering limited insight into temporal understanding over long contexts. In this work, we define a scene as a coherent segment of a video in which both visual and semantic contexts remain consistent, aligning with human perception. This leads us to a key question: can current VLMs reason effectively over long, scene-level contexts? To answer this, we introduce a new benchmark, SceneBench, designed to provide scene-level challenges. Our evaluation reveals a sharp drop in accuracy when VLMs attempt to answer scene-level questions, indicating significant forgetting of long-range context. To further validate these findings, we propose Scene Retrieval-Augmented Generation (Scene-RAG), which constructs a dynamic scene memory by retrieving and integrating relevant context across scenes. This Scene-RAG improves VLM performance by +2.50%, confirming that current models still struggle with long-context retention. We hope SceneBench will encourage future research toward VLMs with more robust, human-like video comprehension.

视频理解长视频上下文记忆基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。