arXiv:2608.27881cs.CV2026-08

通过自演化记忆重构历史数据,提升视觉语言模型的流式视频理解能力

StreamEMS: Streaming Video Understanding with Self-Evolving Memory Scheme for Vision-Language Models

论文配图:StreamEMS: Streaming Video Understanding with Self-Evolving Memory Scheme for Vision-Language Models
图 1 · 摘自论文原文
  • 用语义与先验双重演化机制优化记忆表示
  • 在OVO-Bench和StreamingBench上超越现有方法
  • 高令牌丢失率下仍保持稳定性能,适合资源受限场景

近期许多流式视频理解方法通过构建外部记忆来存储历史数据以降低计算量。多数方法关注当前数据注入(写)和历史信息检索(读),却忽略了提升记忆本身表征能力的机会。本文提出StreamEMS,一种通用机制,通过自演化记忆方案重构内存中的历史数据,实现更丰富、更鲁棒的记忆表征。具体地,引入语义演化模块,通过从粗到细的逐步缩小语义尺度,挖掘有信息量的记忆单元,提升表征密度;同时引入先验感知演化模块,利用历史记忆分布优化当前状态,增强鲁棒性。在OVO-Bench和StreamingBench两个主流数据集上的验证表明,该方法优于现有方法,且在高令牌丢失率设置下优势依然显著,证明了其在释放记忆潜力方面的有效性与鲁棒性。

原文摘要 · Abstract (English)

Recently, many streaming video understanding methods have been proposed by constructing an external memory to store historical data for computational reduction. Most methods focus on optimizing the injection procedure of current data (write) and retrieving informative historical data (read) from memory, while overlooking the opportunity to further enhancing the representational capability of memory itself. In this work, we present StreamEMS, a general mechanism for improving streaming video understanding by re-structuring the historical data stored in memory through self-evolving memory scheme, enabling more informative and robust memory representations. Specifically, we first introduce a Semantic Evolution Module to evolve the memory into more information-dense representations by exploiting informative memory entities discovered via progressively shrinking semantic scales from coarse to fine. In addition, we further introduce a Prior-informed Evolution Module to evolve memory into more robust representations by leveraging prior memory distributions to refine the current memory state. We validate the effectiveness of our proposed designs on widely-used streaming video understanding datasets, i.e., OVO-Bench and StreamingBench, and the results showcase that our method performs better than other methods. Moreover, the advantage of our method becomes consistently evident even under high token usage drop rate settings, indicating the effectiveness and robustness of our method in unleashing the potential of the memory itself.

流式视频记忆机制视觉语言模型自演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。