arXiv:2507.12821cs.AIcs.LG2025-07被引 19

提出新型游戏基准,评估AI在新环境中的快速世界模型构建能力。

Assessing Adaptive World Models in Machines with Novel Games

  • 用精心设计的持续更新结构的新游戏,测试AI通过交互学习环境模型的能力。
  • 强调动态环境下的快速适应性,而非静态数据训练的表征学习。
  • 适合关注通用智能、自适应系统与认知启发式方法的研究者。

人类智能展现出在陌生情境中迅速适应和有效解决问题的惊人能力。我们认为这种深刻适应性本质上与高效构建和优化环境内部表征(即世界模型)密切相关,这一机制称为世界模型归纳。然而,当前人工智能对世界模型的理解与评估仍局限于从海量数据中学习的静态表征,忽视了在新环境中通过互动与探索学习表征的效率与效果。本文结合数十年认知科学成果,提出世界模型归纳视角,并呼吁建立新的评估框架。具体而言,我们提出基于一类具有真实、深层且持续更新结构的‘新游戏’的基准范式,以明确挑战并评估智能体快速构建世界模型的能力。我们希望该框架能推动未来世界模型评估研究,为实现类人快速适应与鲁棒泛化的人工通用智能迈出关键一步。

原文摘要 · Abstract (English)

Human intelligence exhibits a remarkable capacity for rapid adaptation and effective problem-solving in novel and unfamiliar contexts. We argue that this profound adaptability is fundamentally linked to the efficient construction and refinement of internal representations of the environment, commonly referred to as world models, and we refer to this adaptation mechanism as world model induction. However, current understanding and evaluation of world models in artificial intelligence (AI) remains narrow, often focusing on static representations learned from training on massive corpora of data, instead of the efficiency and efficacy in learning these representations through interaction and exploration within a novel environment. In this Perspective, we provide a view of world model induction drawing on decades of research in cognitive science on how humans learn and adapt so efficiently; we then call for a new evaluation framework for assessing adaptive world models in AI. Concretely, we propose a new benchmarking paradigm based on suites of carefully designed games with genuine, deep and continually refreshing novelty in the underlying game structures -- we refer to this class of games as novel games. We detail key desiderata for constructing these games and propose appropriate metrics to explicitly challenge and evaluate the agent's ability for rapid world model induction. We hope that this new evaluation framework will inspire future evaluation efforts on world models in AI and provide a crucial step towards developing AI systems capable of human-like rapid adaptation and robust generalization -- a critical component of artificial general intelligence.

世界模型自适应评估基准通用智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。