arXiv:2604.22409cs.CV2026-04被引 5

评测智能体在动态环境中的空间记忆能力,发现视觉记忆远弱于文字记账。

SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception-Memory Integration in Embodied Environments

论文配图:SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception-Memory Integration in Embodied Environments
图 1 · 摘自论文原文
  • 用动作控制的场景变化构建长时程空间推理测试
  • 模型在无文本辅助时空间记忆严重退化,仅靠视觉输入难维持一致认知
  • 适合研究视觉-语言模型长期记忆机制的学者使用

多模态大模型在静态视觉空间推理上取得进展,但在需要持续根据自身视角更新信念的具身环境中,难以保持长时程空间一致性。本文提出SpaMEM(基于动作序列的空间记忆),一个大规模诊断基准,通过动作触发的场景变换(生成、放置、移除)来隔离空间信念演化机制。该基准基于物理真实数据集,包含1,000个程序生成房屋中超过25,000次交互序列,涵盖10,601,392张高保真图像,覆盖RGB、深度、实例分割和语义分割四种模态。我们构建了三级推理层次共15项诊断任务:一级评估单次观测下的基础空间感知;二级引入文本状态历史以剥离感知噪声,测试时间推理;三级要求从原始视觉流中端到端维持信念,对应真实场景。同时评估短步更新与长程回溯。对多个开源VLM的测试显示,坐标一致定位仍是瓶颈,且从二级到三级性能骤降,暴露模型严重依赖符号记账,缺乏稳健的视觉记忆能力。SpaMEM提供细粒度诊断标准,推动状态表征、信念更新与长程整合机制研究。部分数据已公开于https://huggingface.co/datasets/mill-ct-liao/SpaMEM。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have advanced static visual--spatial reasoning, yet they often fail to preserve long-horizon spatial coherence in embodied settings where beliefs must be continuously revised from egocentric observations under environmental change. We introduce SpaMEM (Spatial Memory from Action Sequences), a large-scale diagnostic benchmark that isolates the mechanics of spatial belief evolution via action-conditioned scene transformations (spawn, place, remove) over long interaction horizons. SpaMEM is built on a physically grounded dataset with 10,601,392 high-fidelity images across four modalities (RGB, depth, instance, semantic segmentation), collected from 25,000+ interaction sequences in 1,000 procedurally generated houses. We formalize embodied spatial reasoning as a three-level hierarchy with 15 diagnostic tasks: Level 1 measures atomic spatial perception from single observations; Level 2 probes temporal reasoning with oracle textual state histories to factor out perceptual noise; and Level 3 requires end-to-end belief maintenance from raw visual streams under the same task dimensions. We further evaluate both short-term (step-wise) updates and long-term (episodic) reconstruction. Benchmarking representative open-source VLM families reveals a consistent stacked bottleneck: coordinate-consistent grounding remains a hard ceiling, and the sharp collapse from Level 2 to Level 3 exposes a pronounced symbolic scaffolding dependency, where models succeed with text-based bookkeeping but struggle to sustain robust visual memory. SpaMEM provides a granular diagnostic standard and motivates explicit mechanisms for state representation, belief revision, and long-horizon episodic integration. A subset of SpaMEM is publicly available at https://huggingface.co/datasets/mill-ct-liao/SpaMEM.

空间推理视觉记忆具身智能多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。