用超图建模视频多对象关系,提升复杂场景理解能力
HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation
- 构建统一超图融合空间关系与因果演化过程
- 在190万帧数据上五项任务均超越现有方法
- 适合视频理解、智能推理等研究者参考
多模态大模型虽推进了视觉语言任务,但在视频场景理解方面仍存不足。为此,视频场景图生成(VidSGG)应运而生,旨在捕捉跨帧多对象关系。然而,现有方法依赖成对连接,难以处理复杂多对象交互与推理。本文提出基于场景超图的多模态大模型(HyperGLM),通过融合实体场景图(捕捉物体间空间关系)与过程图(建模其因果演变),构建统一超图,实现对多向关系与高阶关系的建模。关键创新在于将该超图注入大模型以支持推理。此外,我们构建了新数据集VSGR,包含来自第三人称、第一人称和无人机视角的190万帧视频,支持场景图生成、场景图预测、视频问答、视频描述与关系推理共五项任务。实验表明,HyperGLM在全部五项任务中持续优于当前最佳方法,有效建模并推理多样视频场景中的复杂关系。
原文摘要 · Abstract (English)
Multimodal LLMs have advanced vision-language tasks but still struggle with understanding video scenes. To bridge this gap, Video Scene Graph Generation (VidSGG) has emerged to capture multi-object relationships across video frames. However, prior methods rely on pairwise connections, limiting their ability to handle complex multi-object interactions and reasoning. To this end, we propose Multimodal LLMs on a Scene HyperGraph (HyperGLM), promoting reasoning about multi-way interactions and higher-order relationships. Our approach uniquely integrates entity scene graphs, which capture spatial relationships between objects, with a procedural graph that models their causal transitions, forming a unified HyperGraph. Significantly, HyperGLM enables reasoning by injecting this unified HyperGraph into LLMs. Additionally, we introduce a new Video Scene Graph Reasoning (VSGR) dataset featuring 1.9M frames from third-person, egocentric, and drone views and supports five tasks: Scene Graph Generation, Scene Graph Anticipation, Video Question Answering, Video Captioning, and Relation Reasoning. Empirically, HyperGLM consistently outperforms state-of-the-art methods across five tasks, effectively modeling and reasoning complex relationships in diverse video scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。