arXiv:2501.04336cs.CV2025-01CVPR被引 23

构建视频语义图谱,让大模型更好理解长视频中的时空关系。

Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs

  • 用环境结构化图谱整合关键视频片段,支持空间与时间关联分析。
  • 在多个基准上显著提升模型对时空关系和布局的理解能力。
  • 适合需要深度理解长视频场景的智能系统开发者使用。

长视频理解面临挑战:大型视觉语言模型需在有限上下文窗口内分析时间分散但空间集中的关键时刻。本文提出受‘记忆宫殿’启发的VideoMindPalace框架,将关键视频片段组织为拓扑结构化的语义图谱。该框架通过(i)手物跟踪与交互、(ii)活动区域聚类(代表重复活动的空间区域)、(iii)环境布局映射,使大语言模型能基于真实场景进行自然语言解析,实现对时空及三维上下文的精准理解。此外,我们构建了视频记忆宫殿基准(VMB),用于评估人类级推理能力,包括空间定位、时间推理和布局感知的序列理解。在VMB及多个主流视频问答数据集(EgoSchema、NExT-QA、IntentQA、Active Memories Benchmark)上的实验表明,VideoMindPalace在时空连贯性和人类对齐推理方面取得显著提升,推动了视觉语言模型在长视频分析中的能力边界。

原文摘要 · Abstract (English)

Long-form video understanding with Large Vision Language Models is challenged by the need to analyze temporally dispersed yet spatially concentrated key moments within limited context windows. In this work, we introduce VideoMindPalace, a new framework inspired by the "Mind Palace", which organizes critical video moments into a topologically structured semantic graph. VideoMindPalace organizes key information through (i) hand-object tracking and interaction, (ii) clustered activity zones representing specific areas of recurring activities, and (iii) environment layout mapping, allowing natural language parsing by LLMs to provide grounded insights on spatio-temporal and 3D context. In addition, we propose the Video MindPalace Benchmark (VMB), to assess human-like reasoning, including spatial localization, temporal reasoning, and layout-aware sequential understanding. Evaluated on VMB and established video QA datasets, including EgoSchema, NExT-QA, IntentQA, and the Active Memories Benchmark, VideoMindPalace demonstrates notable gains in spatio-temporal coherence and human-aligned reasoning, advancing long-form video analysis capabilities in VLMs.

长视频理解语义图谱视觉语言模型时空推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。