arXiv:2411.18042cs.CV2024-11CVPR被引 27

用超图建模视频多对象关系,提升复杂场景理解能力

HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation

  • 构建统一超图融合空间关系与因果演化过程
  • 在190万帧数据上五项任务均超越现有方法
  • 适合视频理解、智能推理等研究者参考

多模态大模型虽推进了视觉语言任务,但在视频场景理解方面仍存不足。为此,视频场景图生成(VidSGG)应运而生,旨在捕捉跨帧多对象关系。然而,现有方法依赖成对连接,难以处理复杂多对象交互与推理。本文提出基于场景超图的多模态大模型(HyperGLM),通过融合实体场景图(捕捉物体间空间关系)与过程图(建模其因果演变),构建统一超图,实现对多向关系与高阶关系的建模。关键创新在于将该超图注入大模型以支持推理。此外,我们构建了新数据集VSGR,包含来自第三人称、第一人称和无人机视角的190万帧视频,支持场景图生成、场景图预测、视频问答、视频描述与关系推理共五项任务。实验表明,HyperGLM在全部五项任务中持续优于当前最佳方法,有效建模并推理多样视频场景中的复杂关系。

原文摘要 · Abstract (English)

Multimodal LLMs have advanced vision-language tasks but still struggle with understanding video scenes. To bridge this gap, Video Scene Graph Generation (VidSGG) has emerged to capture multi-object relationships across video frames. However, prior methods rely on pairwise connections, limiting their ability to handle complex multi-object interactions and reasoning. To this end, we propose Multimodal LLMs on a Scene HyperGraph (HyperGLM), promoting reasoning about multi-way interactions and higher-order relationships. Our approach uniquely integrates entity scene graphs, which capture spatial relationships between objects, with a procedural graph that models their causal transitions, forming a unified HyperGraph. Significantly, HyperGLM enables reasoning by injecting this unified HyperGraph into LLMs. Additionally, we introduce a new Video Scene Graph Reasoning (VSGR) dataset featuring 1.9M frames from third-person, egocentric, and drone views and supports five tasks: Scene Graph Generation, Scene Graph Anticipation, Video Question Answering, Video Captioning, and Relation Reasoning. Empirically, HyperGLM consistently outperforms state-of-the-art methods across five tasks, effectively modeling and reasoning complex relationships in diverse video scenes.

视频理解超图建模多模态场景图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。