arXiv:2603.21493cs.CVcs.MM2026-03被引 2

构建统一评估框架,测试视频大模型在真实资源下的实时理解能力

StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding

论文配图:StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding
图 1 · 摘自论文原文
  • 用固定容量内存库标准化历史视觉上下文,统一评估标准
  • 发现现有视频大模型在效率、存储与准确率间差距显著
  • 适合关注视频实时推理部署的研究者和开发者

实时连续的视觉信号理解对现实交互式AI应用至关重要,但当前研究多聚焦于有限视觉上下文下的问答准确率或编码效率提升,忽视了实际部署中资源受限的挑战。为此,我们提出StreamingEval,一个面向视频大模型在真实约束下流式视频理解能力的统一评估框架。该框架在标准化协议下对比主流离线模型与新兴在线视频模型,明确刻画效率、存储与准确率之间的权衡关系。具体而言,采用固定容量的内存银行来规范可访问的历史视觉上下文,并联合评估视觉编码效率、文本解码延迟与任务性能,以量化整体系统可部署性。在多个数据集上的大量实验揭示当前视频大模型与真实流式应用需求之间存在显著差距,为未来研究提供了系统性基准。代码将发布于https://github.com/wwgTang-111/StreamingEval1。

原文摘要 · Abstract (English)

Real-time, continuous understanding of visual signals is essential for real-world interactive AI applications, and poses a fundamental system-level challenge. Existing research on streaming video understanding, however, typically focuses on isolated aspects such as question-answering accuracy under limited visual context or improvements in encoding efficiency, while largely overlooking practical deployability under realistic resource constraints. To bridge this gap, we introduce StreamingEval, a unified evaluation framework for assessing the streaming video understanding capabilities of Video-LLMs under realistic constraints. StreamingEval benchmarks both mainstream offline models and recent online video models under a standardized protocol, explicitly characterizing the trade-off between efficiency, storage and accuracy. Specifically, we adopt a fixed-capacity memory bank to normalize accessible historical visual context, and jointly evaluate visual encoding efficiency, text decoding latency, and task performance to quantify overall system deployability. Extensive experiments across multiple datasets reveal substantial gaps between current Video-LLMs and the requirements of realistic streaming applications, providing a systematic basis for future research in this direction. Codes will be released at https://github.com/wwgTang-111/StreamingEval1.

视频理解流式处理评估框架视频大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。