arXiv:2603.19054cs.CVcs.AI2026-03被引 4

提出新框架,让视频理解模型更高效地主动响应用户查询。

Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding

论文配图:Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
图 1 · 摘自论文原文
  • 将语义理解与视频感知分离,分步处理提升效率
  • 在StreamingBench和OVO-Bench上准确率与效率双提升
  • 适合资源受限场景下的实时视频交互应用

流式视频理解的最新进展催生了模型主动响应用户查询的新范式。现有主动式VideoLLM依赖逐帧触发决策,面临效率与准确率的权衡。我们提出Em-Garde,一种将语义理解与流式感知解耦的新框架。查询时,指令引导的提案解析器将用户查询转化为结构化、感知基础的视觉提案;流式处理阶段,轻量级提案匹配模块通过嵌入匹配高效触发响应。在StreamingBench和OVO-Bench上的实验表明,该框架在主动响应准确率与效率上均优于先前模型,验证了其在严格计算约束下实现有效主动视频理解的可行性。

原文摘要 · Abstract (English)

Recent advances in Streaming Video Understanding has enabled a new interaction paradigm where models respond proactively to user queries. Current proactive VideoLLMs rely on per-frame triggering decision making, which suffers from an efficiency-accuracy dilemma. We propose Em-Garde, a novel framework that decouples semantic understanding from streaming perception. At query time, the Instruction-Guided Proposal Parser transforms user queries into structured, perceptually grounded visual proposals; during streaming, a Lightweight Proposal Matching Module performs efficient embedding-based matching to trigger responses. Experiments on StreamingBench and OVO-Bench demonstrate consistent improvements over prior models in proactive response accuracy and efficiency, validating an effective solution for proactive video understanding under strict computational constraints.

视频理解主动响应流式处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。