评测多模态大模型在长时间对话中的记忆能力,发现现有模型记忆管理仍有短板。
Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
- 构建跨会话多模态对话数据集,含视觉与文本信息
- 13个系统测试显示模型记忆推理与管理能力普遍不足
- 适合研究多模态记忆、对话智能的学者和开发者
长期记忆是多模态大语言模型(MLLM)代理在对话场景中的一项关键能力,但现有基准要么仅评估纯文本多会话记忆,要么只关注局部上下文下的多模态理解,无法衡量多模态记忆在长时间对话轨迹中的保存、组织与演化。为此,我们提出Mem-Gallery,一个用于评估MLLM代理多模态长期对话记忆的新基准。该基准包含高质量的多会话对话,基于视觉与文本信息,具有长交互周期和丰富的多模态依赖关系。基于此数据集,我们构建了系统性评估框架,从记忆提取与运行时适配、记忆推理、记忆知识管理三个功能维度评估关键记忆能力。对13种记忆系统的广泛评测揭示若干关键发现:显式保留多模态信息与组织机制至关重要,记忆推理与知识管理仍存在持续局限,且当前模型存在效率瓶颈。
原文摘要 · Abstract (English)
Long-term memory is a critical capability for multimodal large language model (MLLM) agents, particularly in conversational settings where information accumulates and evolves over time. However, existing benchmarks either evaluate multi-session memory in text-only conversations or assess multimodal understanding within localized contexts, failing to evaluate how multimodal memory is preserved, organized, and evolved across long-term conversational trajectories. Thus, we introduce Mem-Gallery, a new benchmark for evaluating multimodal long-term conversational memory in MLLM agents. Mem-Gallery features high-quality multi-session conversations grounded in both visual and textual information, with long interaction horizons and rich multimodal dependencies. Building on this dataset, we propose a systematic evaluation framework that assesses key memory capabilities along three functional dimensions: memory extraction and test-time adaptation, memory reasoning, and memory knowledge management. Extensive benchmarking across thirteen memory systems reveals several key findings, highlighting the necessity of explicit multimodal information retention and memory organization, the persistent limitations in memory reasoning and knowledge management, as well as the efficiency bottleneck of current models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。