评测多模态智能体记忆的视觉保留能力,发现现有模型难保细节与状态变化。
MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory

- 从场景到像素级证据粒度,评估记忆对视觉信息的保留程度。
- 在8个生活任务中,13种方法均难以应对需追踪视觉状态变化的难题。
- 适合关注视觉记忆、长期推理与模型可解释性的研究者使用。
长时记忆日益呈现多模态特性,但现有评估很少检验智能体是否保留了后续推理所需的视觉证据。以往工作中,许多基于视觉的问题仅靠标题或文本线索即可回答,无需保存细粒度视觉信息。而需要推理视觉状态演变的复杂问题则几乎缺失。为此,我们提出 MemEye 框架,从两个维度评估记忆能力:一是关键视觉证据的粒度(从场景级到像素级),二是检索证据的使用方式(从单一证据到演化合成)。在此框架下,构建涵盖8个生活场景任务的新基准,通过消融验证门控机制评估答案可得性、捷径抗性、视觉必要性及推理结构。在4个视觉语言模型(VLM)骨干上评估13种记忆方法,结果表明当前架构仍难以保留细粒度视觉细节并推理时间状态变化。研究发现,长时多模态记忆依赖于证据路由、时序跟踪与细节提取。
原文摘要 · Abstract (English)
Long-term agent memory is increasingly multimodal, yet existing evaluations rarely test whether agents preserve the visual evidence needed for later reasoning. In prior work, many visually grounded questions can be answered using only captions or textual traces, allowing answers to be inferred without preserving the fine-grained visual evidence. Meanwhile, harder cases that require reasoning over changing visual states are largely absent. Therefore, we introduce MemEye, a framework that evaluates memory capabilities from two dimensions: one measures the granularity of decisive visual evidence (from scene-level to pixel-level evidence), and the other measures how retrieved evidence must be used (from single evidence to evolutionary synthesis). Under this framework, we construct a new benchmark across 8 life-scenario tasks, with ablation-driven validation gates for assessing answerability, shortcut resistance, visual necessity, and reasoning structure. By evaluating 13 memory methods across 4 VLM backbones, we show that current architectures still struggle to preserve fine-grained visual details and reason about state changes over time. Our findings show that long-term multimodal memory depends on evidence routing, temporal tracking, and detail extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。