arXiv:2512.03627cs.AI2025-12被引 25

让AI agent像人一样长期记忆,支持多模态持续学习。

MemVerse: Multimodal Memory for Lifelong Learning Agents

论文配图:MemVerse: Multimodal Memory for Lifelong Learning Agents
图 1 · 摘自论文原文
  • 用分层知识图谱结构化存储多模态长期记忆
  • 通过周期性知识蒸馏实现快速可微召回
  • 适合需要长期记忆的交互式AI系统

尽管大规模语言和视觉模型发展迅速,智能体仍面临根本局限:缺乏可靠记忆。无记忆导致灾难性遗忘、长程推理困难,在多模态或交互环境中表现失序。我们提出MemVerse,一种模型无关、即插即用的记忆框架,融合快速参数化回忆与分层检索式记忆,实现可扩展、自适应的多模态智能。MemVerse维持短期记忆以保存近期上下文,同时将原始多模态经验转化为分层知识图谱形式的长期记忆。该设计支持持续整合、自适应遗忘与有界内存增长。为应对实时需求,引入周期性知识蒸馏机制,将长期记忆中的关键知识压缩至参数模型中,实现快速、可微召回并保持可解释性。大量实验表明,MemVerse显著提升多模态推理与持续学习效率,使智能体在长时间交互中具备记忆、适应与连贯推理能力。

原文摘要 · Abstract (English)

Despite rapid progress in large-scale language and vision models, AI agents still suffer from a fundamental limitation: they cannot remember. Without reliable memory, agents catastrophically forget past experiences, struggle with long-horizon reasoning, and fail to operate coherently in multimodal or interactive environments. We introduce MemVerse, a model-agnostic, plug-and-play memory framework that bridges fast parametric recall with hierarchical retrieval-based memory, enabling scalable and adaptive multimodal intelligence. MemVerse maintains short-term memory for recent context while transforming raw multimodal experiences into structured long-term memories organized as hierarchical knowledge graphs. This design supports continual consolidation, adaptive forgetting, and bounded memory growth. To handle real-time demands, MemVerse introduces a periodic distillation mechanism that compresses essential knowledge from long-term memory into the parametric model, allowing fast, differentiable recall while preserving interpretability. Extensive experiments demonstrate that MemVerse significantly improves multimodal reasoning and continual learning efficiency, empowering agents to remember, adapt, and reason coherently across extended interactions.

长期记忆多模态持续学习知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。