arXiv:2601.16690cs.CLcs.CV2026-01被引 7

用互动游戏评测视觉语言模型的长期记忆能力,发现空间推理仍是难点。

EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents

  • 基于环境轨迹生成可验证问题,覆盖多类记忆技能。
  • 视觉场景中归纳与空间推理表现差,模型提升不明显。
  • 适合研究长时记忆、多模态推理的AI开发者使用。

我们提出EMemBench,一个程序化基准生成器,通过交互式游戏评估智能体的长期情景记忆。不同于固定问题集,该框架从环境驱动的轨迹中生成问题,涵盖纯文本与视觉游戏环境。每个模板基于游戏底层信号计算可验证的真值,保证答案可解且覆盖单跳/多跳回忆、归纳、时间、空间、逻辑及对抗性记忆任务。我们以强语言/视觉语言模型为骨干,使用上下文提示作为基线,在15个文本游戏和多个视觉种子上评估。结果远未饱和:归纳与空间推理在视觉环境中仍为瓶颈。对于开放骨架模型,长期记忆在文本游戏中带来显著提升,但对视觉语言模型效果不一致,表明视觉情境下的情景记忆仍是开放挑战。人类实验进一步揭示了该基准的难度与可解释性。

原文摘要 · Abstract (English)

We introduce EMemBench, a programmatic benchmark generator for evaluating long-term episodic memory of agents through interactive games. Rather than using a fixed set of questions, EMemBench generates questions from environment-grounded trajectories, covering both text-only and visual game environments. Each template computes verifiable ground truth from underlying game signals, with controlled answerability and balanced coverage over memory skills: single/multi-hop recall, induction, temporal, spatial, logical, and adversarial. We evaluate memory agents with strong LMs/VLMs as backbones, using in-context prompting as baselines. Across 15 text games and multiple visual seeds, results are far from saturated: induction and spatial reasoning are persistent bottlenecks, especially in visual settings. Persistent memory yields clear gains for open backbones on text games, but improvements are less consistent for VLM agents, suggesting that visually grounded episodic memory remains an open challenge. A human study further contextualizes the difficulty and interpretability of EMemBench.

记忆评测视觉语言模型长时记忆交互评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。