arXiv:2608.05703cs.CV2026-08

构建小时级视频理解新基准,推动智能体持续交互与长时记忆能力发展。

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

论文配图:StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding
图 1 · 摘自论文原文
  • 提出双层架构StreamMind,分离实时交互与长期记忆处理
  • 在3,646个开放问答上验证,现有方法难以兼顾长时记忆与细节保留
  • 适合研究长时视觉理解、多模态智能体的开发者和研究人员

在连续真实环境中部署自主多模态智能体,需处理无界音视频流并维持小时级记忆。但现有评估多依赖短片段和选择题,导致仅分析最后四帧的简单基线也能超越复杂流模型,且选项暴露语言捷径。我们提出StreamArena,一个面向小时级、交互式流视频理解的基准,包含243段平均88.8分钟的完整视频和3,646个严格标注的开放问答对,评估实时感知、历史回溯、主动交互及多模态工具使用。跨多种系统评估揭示连续交互与长时多模态理解间的矛盾:仅保留近期帧无法恢复远期事件,将过往观察转为文本会丢失视觉证据,反复压缩视觉记忆则难以保持细粒度信息。为此,我们设计StreamMind,通过独立调度的前端任务处理低延迟交互与主动监控,后端异步构建持久化多模态记忆并执行历史召回与外部搜索。StreamMind在四项能力上均优于现有基线,并通过复用持久状态降低查询响应延迟。

原文摘要 · Abstract (English)

Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, current evaluations predominantly rely on brief clips and multiple-choice formats. This design allows minimal baselines that process only the last four frames to match or surpass complex streaming models, while answer options also expose language shortcuts. We introduce StreamArena, a benchmark for hour-scale, interactive streaming video understanding. StreamArena contains 243 full-length videos averaging 88.8 minutes and 3,646 rigorously annotated, open-ended question-answer pairs that evaluate real-time perception, historical retrospection, proactive interaction, and multimodal tool utilization. Evaluation across diverse systems exposes a tension between continuous interaction and long-horizon multimodal comprehension. Methods that retain only recent frames cannot recover distant events, methods that convert past observations into text lose visual evidence, and methods that repeatedly compress visual memory struggle to preserve fine-grained details over time. We address this tension with StreamMind, a two-tier architecture that assigns latency-critical interaction and proactive monitoring to independently scheduled frontend workers, while backend workers asynchronously construct persistent multimodal memory and perform historical recall and external search. StreamMind outperforms existing streaming baselines across all four capabilities and reduces query-to-answer latency by reusing persistent state.

视频理解长时记忆多模态智能体交互式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。