首个面向具身场景的流式视频理解评测基准,测试模型持续感知与推理能力。
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
- 构建分层级的具身任务(感知-交互-规划)和时序推理模式(回溯-实时-前瞻)
- 基于156段长视频生成2.1万条带时间戳的问答对,覆盖42项任务
- 揭示主流视觉大模型在流式具身理解上仍存在显著短板,适合研究具身智能的学者
随着具身智能向真实世界部署推进,持续感知并推理流式视觉输入的能力变得至关重要。在此类场景中,智能体需保持环境态势感知,理解与周围实体的互动,并基于过往观察、当前上下文和未来预期动态规划行为。为推动该方向发展,我们提出首个面向具身场景的流式视频问答评测基准StreamEQA。该基准从具身性和流式性两个正交维度评估现有多模态大语言模型(MLLMs)。在具身维度上,问题分为三类:感知、交互与规划,逐步考察模型识别细粒度视觉细节、推理人机交互以及执行目标导向高阶推理的能力;在流式维度上,问题分为回溯、实时与前瞻推理三种模式,分别依赖不同的时间上下文。StreamEQA基于156段独立长视频构建,定义42项任务,通过自动化生成与人工精修相结合的混合流程,生成约2.1万组带精确时间戳的问答对。对13个前沿视频-大语言模型的评估表明,尽管这些模型在传统基准上表现优异,但在具身场景下的流式视频理解仍面临挑战。我们希望StreamEQA能推动具身应用中流式视频理解的研究进展。
原文摘要 · Abstract (English)
As embodied intelligence advances toward real-world deployment, the ability to continuously perceive and reason over streaming visual inputs becomes essential. In such settings, an agent must maintain situational awareness of its environment, comprehend the interactions with surrounding entities, and dynamically plan actions informed by past observations, current contexts, and anticipated future events. To facilitate progress in this direction, we introduce StreamEQA, the first benchmark designed for streaming video question answering in embodied scenarios. StreamEQA evaluates existing MLLMs along two orthogonal dimensions: Embodied and Streaming. Along the embodied dimension, we categorize the questions into three levels: perception, interaction, and planning, which progressively assess a model's ability to recognize fine-grained visual details, reason about agent-object interactions, and perform high-level goal-directed reasoning. For the streaming dimension, questions are divided into backward, real-time, and forward reasoning, with each mode relying on a distinct temporal context. Built upon 156 independent long videos, StreamEQA defines 42 tasks and generates approximately 21K question-answer pairs with precise timestamps through a hybrid pipeline combining automated generation and human refinement. Evaluations of 13 state-of-the-art video-LLMs reveal that, despite strong performance on conventional benchmarks, these models still struggle with streaming video understanding in embodied scenarios. We hope StreamEQA will catalyze research on streaming video understanding for embodied applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。