只点一下,自动构建视频场景关系图
Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click

- 用户点击后,自动追踪主体并发现互动对象
- 在OpenPVSG上实现高精度时序场景图生成
- 适合需要交互式视频理解的智能系统开发者
当前最先进的视频场景图生成系统为结构化视觉理解提供支持,但作为封闭的前馈流程,无法融入人工指导。而可提示分割模型(如SAM2)虽支持精准用户交互,却缺乏语义与关系推理能力。本文提出Click2Graph,首个交互式全景视频场景图生成框架,融合视觉提示与时空语义理解。仅需一次用户输入(点击或框选),即可实现跨时间的主体分割与跟踪,自动发现互动对象,并生成<主体, 客体, 谓词>三元组构成时序一致的场景图。框架包含两个核心组件:动态交互发现模块(生成主体条件下的客体提示)和联合实体-谓词分类头。在OpenPVSG基准测试中,Click2Graph建立了用户引导式全景视频场景图生成的新基线,验证了人类提示与全景定位、关系推理结合的可行性,实现可控且可解释的视频理解。
原文摘要 · Abstract (English)
State-of-the-art Video Scene Graph Generation (VSGG) systems provide structured visual understanding but operate as closed, feed-forward pipelines with no ability to incorporate human guidance. In contrast, promptable segmentation models such as SAM2 enable precise user interaction but lack semantic or relational reasoning. We introduce Click2Graph, the first interactive framework for Panoptic Video Scene Graph Generation (PVSG) that unifies visual prompting with spatial, temporal, and semantic understanding. From a single user cue, such as a click or bounding box, Click2Graph segments and tracks the subject across time, autonomously discovers interacting objects, and predicts <subject, object, predicate> triplets to form a temporally consistent scene graph. Our framework introduces two key components: a Dynamic Interaction Discovery Module that generates subject-conditioned object prompts, and a Semantic Classification Head that performs joint entity and predicate reasoning. Experiments on the OpenPVSG benchmark demonstrate that Click2Graph establishes a strong foundation for user-guided PVSG, showing how human prompting can be combined with panoptic grounding and relational inference to enable controllable and interpretable video scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。