arXiv:2503.14427cs.AI2025-03EMNLP被引 5

评测AI在动态虚拟密室中探索决策能力的基准数据集

VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Rooms

  • 构建20个动态变化的虚拟密室环境,评估模型探索与规划能力
  • 顶尖多模态模型在该基准上普遍失败,解谜进度差异大
  • 引入记忆与推理机制可显著提升探索效率和试错能力

密室逃脱对认知能力提出独特挑战:玩家仅需‘逃出房间’,却必须主动探索环境、收集信息,并通过反复试错找到解决方案。为此,我们提出VisEscape,一个包含20个虚拟密室的基准,专门用于评估人工智能模型在探索驱动型决策中的表现。成功不仅依赖于解决孤立谜题,更需要持续构建和修正时空知识。在VisEscape上,即使最先进的多模态模型也普遍无法逃出房间,表现出显著的进展差异与问题解决方式多样性。研究发现,集成记忆管理和推理机制有助于高效探索,并支持连续假设生成与验证,从而在动态、探索驱动环境中实现显著性能提升。

原文摘要 · Abstract (English)

Escape rooms present a unique cognitive challenge that demands exploration-driven planning: with the sole instruction to 'escape the room', players must actively search their environment, collecting information, and finding solutions through repeated trial and error. Motivated by this, we introduce VisEscape, a benchmark of 20 virtual escape rooms specifically designed to evaluate AI models under these challenging conditions, where success depends not only on solving isolated puzzles but also on iteratively constructing and refining spatial-temporal knowledge of a dynamically changing environment. On VisEscape, we observe that even state-of-the-art multi-modal models generally fail to escape the rooms, showing considerable variation in their progress and problem-solving approaches. We find that integrating memory management and reasoning contributes to efficient exploration and enables successive hypothesis formulation and testing, thereby leading to significant improvements in dynamic and exploration-driven environments

密室逃脱探索决策基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。