arXiv:2607.09759cs.CVcs.AI2026-07

让AI像人一样记住视频中持续出现的实体,实现长期记忆追踪。

ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams

论文配图:ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams
图 1 · 摘自论文原文
  • 以实体为中心构建多尺度记忆,突破传统帧级存储限制。
  • 在六个长视频与终身记忆基准上全部达到最优性能。
  • 可接入现有助手系统,支持真实世界开放视频流处理。

构建能持续观看世界、记住所见并基于积累经验推理的智能体是长期目标。近期具备视频流长期记忆的多模态智能体备受关注。然而,现有系统将记忆存储于模型上下文或扁平特征库中,按帧组织而非围绕持续存在的实体,导致仅限于有限视频,难以追踪随时间重复出现的人或物。本文提出ReflectWorld-MM,一种面向开放视频流的实体导向多模态记忆系统,包含三部分:感知前端将音视频流转化为短时记忆内的实体解析观测;基于人类记忆理论的分层长期记忆,融合多尺度情景记忆、动态演化的实体中心语义记忆和程序性记忆;完整实现系统可处理任意视频流,并无缝集成至现成智能体。在六项长视频与终身记忆基准测试中,ReflectWorld-MM在所有任务上均取得最佳准确率,超越强基线模型与前沿方法。

原文摘要 · Abstract (English)

Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest. Unfortunately, existing systems either keep their memory inside the model context or in a flat feature store, and organize it around frames rather than around the persistent entities a stream is really about, which confines them to bounded videos and weakens their ability to track who and what reappears over time. In this paper, we propose ReflectWorld-MM, an entity-oriented multimodal memory system for open-ended video streams. It consists of three parts. The first is a perception front-end that turns an audiovisual stream into entity-resolved observations under a bounded short-term memory. The second is a hierarchical long-term memory, grounded in human memory theory, that couples a multi-scale episodic memory, an evolving entity-centric semantic memory, and a procedural memory. The third is a complete realization, built for real-world operation, that ingests arbitrary streams and plugs into off-the-shelf assistants. Across six long-video and lifelong-memory benchmarks, ReflectWorld-MM achieves the best accuracy on all six, outperforming strong memory agents and a frontier model.

多模态记忆视频理解实体追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。