arXiv:2507.05169cs.LGcs.AI2025-07被引 17

提出通用世界模型新架构,支持多层级推理与自主行动。

Critique of World Model

  • 基于状态化分层表示,融合连续与离散建模,构建生成式自监督框架。
  • 设计可模拟真实世界所有可行动态的通用世界模型,支撑目的性推理。
  • 适合研究通用人工智能、具身智能与虚拟代理的学者参考。

世界模型作为生物体所处真实环境的算法模拟器,近年来因发展具备人工智能的虚拟代理需求而兴起。本文从科幻经典《沙丘》中的想象出发,借鉴心理学中的‘假设性思维’概念,主张世界模型的核心目标是为有目的的推理与行动模拟现实世界的全部可行动态。我们分析了世界建模的关键维度:数据、表征、架构、学习目标与使用方式,评估现有方法及其权衡。在此基础上,提出一种通用生成潜变量预测(GLP)架构,采用状态化、分层、多层级、混合连续/离散表征,结合生成式与自监督学习框架,展望由该模型驱动的物理性、代理性和嵌套式(PAN)通用智能系统。

原文摘要 · Abstract (English)

World Model, the algorithmic simulator of the real-world environment which biological agents experience and act upon, has been an emerging topic in recent years due to the rising need to develop virtual agents with artificial (general) intelligence. There has been much discussion on what a world model really is, how to build it, how to use it, and how to evaluate it. In this essay, starting from the imagination in the famed Sci-Fi classic Dune, and drawing inspiration from the concept of ``hypothetical thinking'' in psychology literature, we argue the primary goal of a world model to be {\it simulating all actionable possibilities of the real world for purposeful reasoning and acting}. We examine the key design dimensions of world modeling: data, representation, architecture, learning objective, and usage, surveying existing approaches and analyzing their tradeoffs. Building on this examination, we propose a new Generative Latent Prediction (GLP) architecture for a general-purpose world model, based on stateful, hierarchical, multi-level, and mixed continuous/discrete representations, and a generative and self-supervised learning framework, with an outlook of a Physical, Agentic, and Nested (PAN) AGI system enabled by such a model.

世界模型通用智能生成模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。