arXiv:2606.05008cs.CVcs.AI2026-06

首个面向多模态模型记忆能力的系统评估框架,揭示模型在记忆保持与干扰下的真实表现。

M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks

论文配图:M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks
图 1 · 摘自论文原文
  • 基于认知心理学设计隔离记忆维度的任务,精准探测模型记忆机制
  • 发现模型处理多视频流时难以保持独立表征,且干扰模式与人类差异显著
  • 适合关注多模态模型长期记忆、认知机制研究的研究者

随着多模态模型向长视频理解发展,记忆能力成为关键。尽管已有大量视频数据集和基准,但现有工作主要聚焦感知与推理,缺乏对记忆能力的系统评估:模型保留了什么信息、信息保真度如何、在干扰下记忆是否稳健。为此,我们提出M$^3$Eval,首个全面评估多模态模型记忆维度的框架与基准。其设计基于认知心理学,任务精心构建以分离记忆的关键方面。利用M$^3$Eval,我们在代表性多模态模型上开展广泛实验,揭示出一致的弱点与独特行为:模型在处理并行视频流时难以维持解耦表征;表现出的干扰模式与人类记忆显著不同;更可靠地将记忆锚定在空间域而非时间域;符号记忆能力有限。该基准为未来研究提供宝贵资源,研究结果强调记忆是基础却未被充分探索的能力,并为设计更有效的记忆机制提供洞见。代码与数据集详见 https://pku-value-lab.github.io/m3eval-homepage。

原文摘要 · Abstract (English)

As multi-modal models advance towards long-form video understanding, memory emerges as a critical capability. Despite substantial efforts in developing video datasets and benchmarks, existing works primarily focus on perception and reasoning, without systematically evaluating memory: what models retain, how faithfully information is preserved, and how robust memory remains under interference. To address this gap, we introduce M$^3$Eval, the first comprehensive evaluation framework and benchmark for probing different memory dimensions in multi-modal models. Grounded in cognitive psychology, our design features carefully constructed tasks that isolate key aspects of memory. Leveraging M$^3$Eval, we conduct extensive experiments across representative multi-modal models, revealing consistent weaknesses and distinctive behaviors. We find that models struggle to maintain disentangled representations when processing parallel video streams, exhibit interference patterns differing substantially from those observed in human memory, ground memory sources more reliably in the spatial domain than the temporal domain, and demonstrate limited symbolic memory. Collectively, our benchmark provides a valuable resource for future research, while our findings highlight memory as a fundamental yet underexplored capability and offer insights for designing more effective memory mechanisms in multi-modal models. Our code and dataset are available at https://pku-value-lab.github.io/m3eval-homepage.

多模态记忆评估认知建模视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。