构建游乐场模拟器,评估AI在复杂决策中的长期规划与空间推理能力
Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions
- 设计迷你游乐园仿真环境,整合多类真实决策挑战
- 专家人类表现优于现有智能体11.4倍(简单模式)至15.3倍(中等模式)
- 适合研究长期规划、不确定性建模与自适应决策的学者使用
尽管人工智能发展迅速,但现有系统在真实世界决策中仍面临相互关联的挑战:需要开放式的优化能力,从稀疏经验中主动学习环境动态,于随机环境中进行长周期规划,并对空间信息进行推理。然而,当前的人机基准测试尚未全面评估智能体在具体情境下融合这些能力的表现。为此,我们提出了迷你游乐园(Mini Amusement Parks, MAPs),一个专为评估智能体建模环境、预测不确定性下的长期后果并战略性运营复杂业务而设计的游乐场模拟器。我们提供了专家人类表现数据,并对最先进智能体进行了全面评估,结果显示,在简单模式下专家表现超出现有系统11.4倍,在中等模式下达15.3倍。分析揭示了当前方法在长期规划、样本高效学习、空间推理和不确定性建模方面的持续缺陷。通过将上述挑战统一于单一环境,MAPs为评估具备自适应决策能力的智能体提供了新基准。代码开源:https://github.com/Skyfall-Research/MAPs
原文摘要 · Abstract (English)
Despite rapid progress in artificial intelligence, current systems struggle with the interconnected challenges that define real-world decision making. Practical domains such as business management require open-ended optimization, actively learning environment dynamics from sparse experience, planning over long horizons in stochastic settings, and reasoning over spatial information. Yet no existing human--AI benchmarks assess how well agents integrate these challenges in a grounded decision-making context. To this end, we introduce Mini Amusement Parks (MAPs), an amusement-park simulator designed to evaluate an agent's ability to model its environment, anticipate long-term consequences under uncertainty, and strategically operate a complex business. We provide expert human performance and a comprehensive evaluation of state-of-the-art agents, finding experts outperform these systems by 11.4x on easy mode and 15.3x on medium mode. Our analysis reveals persistent weaknesses in long-horizon planning, sample-efficient learning, spatial reasoning, and modelling uncertainty. By unifying these challenges within a single environment, MAPs offers a new foundation for benchmarking agents capable of adaptable decision making. Code: https://github.com/Skyfall-Research/MAPs
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。