arXiv:2501.15953cs.IRcs.CV2025-01被引 13

用图结构+大模型理解长视频,更准且只看少量帧

Understanding Long Videos via LLM-Powered Entity Relation Graphs

  • 构建动态实体关系图,追踪视频中物体随时间的变化
  • 在EgoSchema和NExT-QA上分别提升2.2和2.0分,平均仅需8.2/8.1帧
  • 适合需要高效理解长视频的场景,如自动剪辑、智能检索

长视频分析在人工智能中面临独特挑战,尤其在跨时间跟踪和理解视觉元素时。现有逐帧处理方法难以维持对象连贯性,尤其当对象短暂消失后再次出现时。其关键局限在于对时间关系的理解不足,难以识别关键片段。为此,我们提出GraphVideoAgent,结合图结构对象追踪与大语言模型能力。系统采用动态图结构,映射并监控视频序列中视觉实体间的演化关系,实现对物体交互与变化的更精细理解,通过全面上下文感知提升帧选择效果。在EgoSchema数据集上的测试显示,相比现有方法提升2.2分,平均仅需8.2帧;在NExT-QA基准上提升2.0分,平均仅需8.1帧。结果表明,该图引导方法在长视频理解任务中显著提升了准确率与计算效率。

原文摘要 · Abstract (English)

The analysis of extended video content poses unique challenges in artificial intelligence, particularly when dealing with the complexity of tracking and understanding visual elements across time. Current methodologies that process video frames sequentially struggle to maintain coherent tracking of objects, especially when these objects temporarily vanish and later reappear in the footage. A critical limitation of these approaches is their inability to effectively identify crucial moments in the video, largely due to their limited grasp of temporal relationships. To overcome these obstacles, we present GraphVideoAgent, a cutting-edge system that leverages the power of graph-based object tracking in conjunction with large language model capabilities. At its core, our framework employs a dynamic graph structure that maps and monitors the evolving relationships between visual entities throughout the video sequence. This innovative approach enables more nuanced understanding of how objects interact and transform over time, facilitating improved frame selection through comprehensive contextual awareness. Our approach demonstrates remarkable effectiveness when tested against industry benchmarks. In evaluations on the EgoSchema dataset, GraphVideoAgent achieved a 2.2 improvement over existing methods while requiring analysis of only 8.2 frames on average. Similarly, testing on the NExT-QA benchmark yielded a 2.0 performance increase with an average frame requirement of 8.1. These results underscore the efficiency of our graph-guided methodology in enhancing both accuracy and computational performance in long-form video understanding tasks.

长视频理解图神经网络大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。