arXiv:2604.18459cs.CVcs.AI2026-04被引 3

让视频模型在线推理时精准响应,且过程透明可追溯。

Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions

论文配图:Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions
图 1 · 摘自论文原文
  • 分离思考与记忆,用可观察指标控制回答时机
  • 在StreamingBench上达71.6%准确率,超越前代方法
  • 适合需要实时决策透明性的智能监控等场景

在真实环境运行的视觉代理必须在视频流中首次出现足够证据时精确响应,而传统离线评估的视频大模型忽视了这一点。转向在线流式处理带来新挑战:决策不透明、响应时间难以对齐视觉证据、在严苛计算预算下维持全局因果一致理解。为此,我们提出一种新框架,将推理控制与记忆整合解耦。其中, extbf{ extmodel{}} 包含两个核心组件:第一,主动思考决策器(ATDM)作为透明推理控制器,通过可观测的进展($oldsymbolρ$)和置信度($oldsymbol{c}$)指标外化决策过程,确保响应时间 $t_r$ 精确匹配首个充分证据时刻 $t^/star$,并实时向用户展示推理链;第二,分层渐进语义融合(HPSI)模块作为高效记忆系统,采用可学习的多层级聚合标记,在片段间传递以构建丰富全局认知状态,且不超限令牌预算。该方法在关键在线视频理解基准上达到新标准,于StreamingBench上取得71.6%、OVOBench上46.9%的准确率。实验表明,Thinking-QwenVL将前代最优在StreamingBench上的准确率从67.63%提升至71.60%。

原文摘要 · Abstract (English)

Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability that is overlooked by conventional video LLMs evaluated in offline settings. The shift to an online, streaming paradigm introduces significant challenges: a lack of decision transparency, the difficulty of aligning response timing with visual evidence, and the need to maintain a global, causally consistent understanding under tight computational budgets. To address these issues, we propose a novel framework that decouples reasoning control from memory integration. We introduce \textbf{\model{}}, an instantiation of this framework with two core components. First, the \emph{Active Thinking Decision Maker (ATDM)} is a transparent reasoning controller that externalizes its decision process using observable progress ($\boldsymbolρ$) and confidence ($\boldsymbol{c}$) metrics. This allows it to precisely time its response $t_r$ to match the first-sufficient-evidence timestamp $t^\star$ while streaming its reasoning to the user. Second, the \emph{Hierarchical Progressive Semantic Integration (HPSI)} module acts as an efficient memory system. It employs a set of learnable, multi-level aggregation tokens that are propagated across clips to build a rich, global cognitive state without exceeding token budgets. %Our approach sets a new standard on key online video understanding benchmarks, achieving strong performance of \textbf{71.6\%} on StreamingBench and \textbf{46.9\%} on OVOBench, demonstrating a robust solution for evidence-aligned and transparent online video analysis. Extensive experiments demonstrate the effectiveness of ATDM and HPSI, e.g., Thinking-QwenVL improves the accuracy of the previous state-of-the-art from 67.63\% to 71.60\% on the StreamingBench benchmark.

在线视频理解透明推理流式处理视频大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。