arXiv:2605.07575cs.CVcs.AI2026-05ACL被引 2

用场景图显式建模视频证据与响应条件,实现更准确的主动视频理解。

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding

论文配图:Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding
图 1 · 摘自论文原文
  • 通过流式视频片段生成查询引导的场景图,建立视觉证据与响应条件的结构化对齐。
  • 在基准测试中,主动与被动任务表现均优于现有方法,提升响应决策准确性。
  • 无需微调,适合需要实时判断何时响应的视频理解应用场景。

主动流式视频理解要求视频大模型在视频播放过程中决定何时响应,现有方法因对视觉证据采用隐式、查询无关的建模而表现不佳。我们提出Response-G1框架,通过场景图建立累积视频证据与查询期望响应条件之间的显式、结构化对齐。该框架包含三个免微调阶段:(1) 从流式片段中在线生成查询引导的场景图;(2) 基于记忆检索最语义相关的历史场景图;(3) 采用检索增强的触发提示,实现每帧“静默/响应”决策。通过共享图表示同时表征证据与条件,Response-G1实现了更可解释且更准确的响应时机判断。在多个基准测试中,该方法在主动与被动任务上均展现出优越性能,验证了显式场景图建模与检索机制在流式视频理解中的优势。

原文摘要 · Abstract (English)

Proactive streaming video understanding requires Video-LLMs to decide when to respond as a video unfolds, a task where existing methods often fall short due to their implicit, query-agnostic modeling of visual evidence. We introduce Response-G1, a novel framework that establishes explicit, structured alignment between the accumulated video evidence and the query's expected response conditions via scene graphs. The framework operates in three fine-tuning-free stages: (1) online query-guided scene graph generation from streaming clips; (2) memory-based retrieval of the most semantically relevant historical scene graphs; and (3) retrieval-augmented trigger prompting for per-frame "silence/response" decisions. By grounding both evidence and conditions in a shared graph representation, Response-G1 achieves more interpretable and accurate response timing decisions. Experimental results on established benchmarks demonstrate the superiority of our method in both proactive and reactive tasks, validating the advantage of explicit scene graph modeling and retrieval in streaming video understanding.

视频理解场景图主动响应流式处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。