用视频编码结构提升长视频理解,兼顾运动细节与效率
GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning

- 通过GOP分组训练运动代理,融合视频编码机制
- 在MotionBench和Egoschema上实现领先VQA性能
- 适合需要精准运动分析的长视频应用开发者
尽管代理式长视频理解已取得进展,现有方法仍缺乏对运动细节的深入理解及高效的记忆架构。本文提出GOPAgen,首次将视频编码器引入视频理解框架,通过在视频编码中的图像组(GOPs)上训练运动代理,实现运动信息的精细建模。进一步设计了与视频编码天然对齐的GOP树推理算法,增强模型对局部运动细节的理解能力。同时,构建了结构化记忆机制,将局部运动信息与结构化页面中的详细描述相结合,并提出粗粒度到细粒度的逐层检索算法以充分挖掘该记忆。此外,引入运动向量数据库,支持不同粒度下的高效运动向量检索。整体方法在多个视频理解基准上表现优异,包括MotionBench和Egoschema,验证了框架的有效性。
原文摘要 · Abstract (English)
Despite significant progress in agentic long video understanding, existing methods still lack detailed motion comprehension coupled with an efficient memory architecture. In this paper, we propose GOPAgen, a novel approach that first integrates video codec into the video understanding framework via a meticulously designed motion agent trained on Groups of Pictures (GOPs) from video codec. We further develop a GOP tree reasoning algorithm, which is naturally aligned with video codec and enhances the model's ability to understand local detailed motions in videos. Additionally, we carefully design a structural memory mechanism that integrates local motion information with detailed captions in structural pages, and propose an efficient coarse-to-fine zoom-in algorithm to fully exploit the structural memory. Furthermore, we incorporate a motion vector database into the framework to enable efficient retrieval of motion vectors at different granularities. Overall, our method achieves superior Video Question Answering (VQA) performance on various video understanding benchmarks, including MotionBench and Egoschema, thereby demonstrating the superiority of our proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。