arXiv:2603.08620cs.CV2026-03中稿 · CVPR被引 10

让视频模型学会在恰当时间回答,避免过早猜测或错过时机。

StreamReady: Learning What to Answer and When in Long Streaming Videos

  • 引入感知时机的答题评分机制,区分何时答对、何时答准。
  • 在多个长视频数据集上超越现有方法,尤其在实时问答任务中表现优异。
  • 适合需要精准响应的场景,如智能客服、实时监控系统。

流式视频理解常涉及时效性场景,模型需在视觉证据出现时即时作答:提前回答属猜测,延迟则丧失实时价值。为此,本文提出答案就绪度评分(ARS),一种带不对称早/晚惩罚的时间敏感目标函数。结合准确率,ARS定义了衡量模型是否适时作答的有效准确率。基于此,我们构建StreamReady框架,通过轻量级就绪机制判断是否已积累足够证据再作答。为评估该能力,我们引入ProReady-QA基准,包含标注的答案证据窗口与跨局部和全局上下文的主动多轮问题。StreamReady在ProReady-QA上表现优异,并在八个额外流式及离线长视频基准上持续领先,展现强大且广泛适用的视频理解能力。

原文摘要 · Abstract (English)

Streaming video understanding often involves time-sensitive scenarios where models need to answer exactly when the supporting visual evidence appears: answering before the evidence reflects speculation, answering after it has passed reduces real-time utility. To capture this behavior, we introduce a readiness-aware formulation of streaming video understanding with the Answer Readiness Score (ARS), a timing-aware objective with asymmetric early and late penalties. When combined with correctness, ARS defines an effective accuracy that measures not just whether a model is right, but whether it answers at the appropriate moment. Building on this formulation, we introduce StreamReady, a framework to unify temporal reasoning with on-time answering through a lightweight readiness mechanism that decides if sufficient evidence has been observed before responding. To evaluate this capability, we further introduce ProReady-QA, a benchmark with annotated answer evidence windows and proactive multi-turn questions across local and global contexts. StreamReady achieves superior performance on ProReady-QA, and consistently outperforms prior methods across eight additional streaming and offline long-video benchmarks, demonstrating robust and broadly generalizable video understanding capability.

视频理解实时问答长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。