新基准测试动态空间推理,挑战模型在变化环境中的长期记忆与规划能力。
EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
- 设计局部可观测的动态迷宫与匹配消除任务,模拟真实环境变化。
- 主流模型在动态场景中表现显著下降,暴露长期记忆短板。
- 引入基于主观体验的记忆机制,支持跨任务经验迁移。
现有空间推理基准多聚焦静态或全局可观测环境,未能捕捉局部感知、环境反馈与全局目标耦合下的长时推理与记忆利用挑战。我们提出两个动态空间基准:局部可观测迷宫导航与匹配-2消除任务,系统评估模型在空间理解与自适应规划方面的能力。每个动作都会引发环境结构变化,要求认知与策略持续更新。我们进一步提出基于主观体验的记忆机制,实现跨任务经验迁移与验证。实验表明,该基准揭示了主流模型在动态空间推理与长期记忆方面的关键缺陷,为未来方法发展提供全面评估平台。代码与数据已公开于 https://anonymous.4open.science/r/EvoEmpirBench-143C/。
原文摘要 · Abstract (English)
Most existing spatial reasoning benchmarks focus on static or globally observable environments, failing to capture the challenges of long-horizon reasoning and memory utilization under partial observability and dynamic changes. We introduce two dynamic spatial benchmarks, locally observable maze navigation and match-2 elimination that systematically evaluate models' abilities in spatial understanding and adaptive planning when local perception, environment feedback, and global objectives are tightly coupled. Each action triggers structural changes in the environment, requiring continuous update of cognition and strategy. We further propose a subjective experience-based memory mechanism for cross-task experience transfer and validation. Experiments show that our benchmarks reveal key limitations of mainstream models in dynamic spatial reasoning and long-term memory, providing a comprehensive platform for future methodological advances. Our code and data are available at https://anonymous.4open.science/r/EvoEmpirBench-143C/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。