系统梳理世界模型的两大功能:理解现实与预测未来。
Understanding World or Predicting Future? A Comprehensive Survey of World Models

- 按理解世界与预测未来划分世界模型类型
- 覆盖生成游戏、自动驾驶等四大应用场景
- 适合关注AGI底层机制的研究者阅读
世界模型因多模态大语言模型(如GPT-4)和视频生成模型(如Sora)的发展而受到广泛关注,是实现通用人工智能的关键。本文对世界模型研究进行系统综述,将其主要功能归纳为两类:一是构建内部表征以理解世界运行机制;二是预测未来状态以实现模拟与决策指导。本文首先梳理这两类模型的最新进展,随后探讨其在生成游戏、自动驾驶、机器人以及社会模拟等关键领域的应用,分析各领域如何利用这些功能。最后,指出当前面临的挑战,并提出未来研究方向。代表性论文及其代码仓库已整理至https://github.com/tsinghua-fib-lab/World-Model。
原文摘要 · Abstract (English)
The concept of world models has garnered significant attention due to advancements in multimodal large language models such as GPT-4 and video generation models such as Sora, which are central to the pursuit of artificial general intelligence. This survey offers a comprehensive review of the literature on world models. Generally, world models are regarded as tools for either understanding the present state of the world or predicting its future dynamics. This review presents a systematic categorization of world models, emphasizing two primary functions: (1) constructing internal representations to understand the mechanisms of the world, and (2) predicting future states to simulate and guide decision-making. Initially, we examine the current progress in these two categories. We then explore the application of world models in key domains, including generative games, autonomous driving, robotics, and social simulacra, with a focus on how each domain utilizes these aspects. Finally, we outline key challenges and provide insights into potential future research directions. We summarize the representative papers along with their code repositories in https://github.com/tsinghua-fib-lab/World-Model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。