arXiv:2605.18577cs.CV2026-05

首个综合评估视频流主动理解能力的基准,涵盖多模态感知与实时响应。

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

论文配图:OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
图 1 · 摘自论文原文
  • 构建覆盖9子任务的多模态视频理解数据集,84%样本需音频信号
  • 提出探针与在线双模式评测,验证模型主动响应与长期鲁棒性
  • 发现语音利用差异大、长期表现退化、非语音音频感知弱

全景主动式视频流理解——即从连续音视频流中自主决定何时发言及说什么——是多模态大语言模型的新兴能力。现有基准在三方面存在不足:主要依赖视觉信号、采用轮询或固定时间戳协议而非真正的主动评估、任务覆盖范围有限,难以可靠评估和区分全景主动式模型。我们提出 OmniPro,首个联合评估多模态感知、主动响应与多样化视频理解任务的基准。它包含2,700个经人工验证的样本,覆盖9个子任务和3个认知层级,涵盖6种基础视频理解能力。其中84%的样本需要音频信号(语音或非语音),每个样本均标注模态隔离标签,支持细粒度多模态分析。我们进一步引入双模式评估协议:探针模式通过在每个真实触发点前后提问来评估内容理解;在线模式要求模型在持续输入中自主决定响应时机,以评估完整主动能力。对11个代表性模型的评估揭示三个关键发现:(1) 音频带来稳定提升,但各模型利用率差异显著;(2) 性能随时间显著下降,表明长时程鲁棒性不足;(3) 非语音音频感知仍是最薄弱环节。

原文摘要 · Abstract (English)

Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-modal large language models. Existing benchmarks fall short in three key aspects: they rely primarily on visual signals, adopt polling or fixed-timestamp protocols instead of true proactive evaluation, and cover only a limited range of tasks, preventing reliable assessment and differentiation of omni-proactive streaming models. We present OmniPro, the first benchmark to jointly evaluate omni-modal perception, proactive responding, and diverse video understanding tasks. It comprises 2,700 human-verified samples spanning 9 sub-tasks and 3 cognitive levels, covering 6 basic video understanding capabilities. Notably, 84% of samples require audio signals (speech or non-speech), and each sample is annotated with modality-isolation labels to enable fine-grained multimodal analysis. We further introduce a dual-mode evaluation protocol: Probe mode assesses content understanding by querying the model before and after each ground-truth trigger, while Online mode evaluates full proactive ability by requiring models to autonomously decide when to respond in streaming input. Evaluating 11 representative models reveals three key findings: (1) audio provides consistent gains but with highly variable utilization across models, (2) performance degrades significantly over time, indicating limited long-horizon robustness, and (3) non-speech audio perception remains the weakest dimension.

视频理解多模态主动响应评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。