arXiv:2511.12027cs.CVcs.AI2025-11被引 4

用结构化记忆增强视频理解,让大模型看长视频更准

GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory

  • 设计图示与叙事式情景记忆,组织事件关系解决长期依赖
  • 在Video-MME长视频测试集上提升23.5%准确率,达73.4%
  • 适合需要深度视频推理的科研与工业应用

长视频理解仍是多模态大模型(MLLMs)的重大挑战,源于固有的令牌限制及捕捉长期时序依赖的复杂性。现有方法难以把握全局上下文与复杂事件关系,影响深层视频推理。为此,我们提出GCAgent——一种全局上下文感知的智能体框架,实现全面的长视频理解。其核心创新是图示与叙事式情景记忆,将事件及其因果和时序关系结构化为简洁有序的上下文,从根本上解决长期依赖问题。GCAgent采用感知-行动-反思多阶段循环,通过记忆管理器检索相关情景上下文,实现鲁棒的上下文感知推理。大量实验表明,GCAgent显著提升长视频理解能力,在Video-MME Long测试集上相较强基线提升23.5%准确率;在同类7B规模模型中达到73.4%的最高长视频准确率,整体平均分71.9%,验证了该智能体推理范式与结构化记忆在类认知长视频理解中的有效性。

原文摘要 · Abstract (English)

Long-video understanding remains a significant challenge for Multimodal Large Language Models (MLLMs) due to inherent token limitations and the complexity of capturing long-term temporal dependencies. Existing methods often fail to capture the global context and complex event relationships necessary for deep video reasoning. To address this, we introduce GCAgent, a novel Global-Context-Aware Agent framework that achieves comprehensive long-video understanding. Our core innovation is the Schematic and Narrative Episodic Memory. This memory structurally models events and their causal and temporal relations into a concise, organized context, fundamentally resolving the long-term dependency problem. Operating in a multi-stage Perception-Action-Reflection cycle, our GCAgent utilizes a Memory Manager to retrieve relevant episodic context for robust, context-aware inference. Extensive experiments confirm that GCAgent significantly enhances long-video understanding, achieving up to 23.5\% accuracy improvement on the Video-MME Long split over a strong MLLM baseline. Furthermore, our framework establishes state-of-the-art performance among comparable 7B-scale MLLMs, achieving 73.4\% accuracy on the Long split and the highest overall average (71.9\%) on the Video-MME benchmark, validating our agent-based reasoning paradigm and structured memory for cognitively-inspired long-video understanding.

视频理解长视频记忆机制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。