arXiv:2608.02713cs.CVcs.AI2026-08

将世界模型转向以智能体为中心,提供更实用的反馈支持持续进化。

Quo Vadis, World Modeling?

论文配图:Quo Vadis, World Modeling?
图 1 · 摘自论文原文
  • 从预测物理状态转为预测智能体可用信息,如执行结果、经验等。
  • 提出六类代理形式与三级赋能层次,系统化设计世界模型。
  • 适合研究持续学习、自主智能体与仿真环境构建的学者。

持续改进的智能体需要动态交互反馈,而直接现实交互成本高、速度慢且难以并行。世界模型提供了一个低成本、可控的中间代理,使智能体在真实行动前获得反馈。传统世界模型主要通过未来物理状态预测实现,但这一范式对需超出原始状态转移的可操作反馈的智能体而言过于局限。本文提出以智能体为中心的交互式世界代理,将核心范式从物理状态转移转向智能体可用的信息转移,如执行结果、检索到的经验或技能、验证信号等,从而拓宽世界模型的应用范围。我们系统性地将世界代理分为六种功能形式:动力学、空间、执行、记忆/经验、技能和奖励/验证代理,它们共同刻画了世界模型支持智能体进化的关键路径。进一步分析这些代理如何在三个渐进层级上赋能智能体:L.1 推理时指导,利用代理输出增强上下文信息以优化决策;L.2 训练时优化,利用代理输出生成奖励、批评或合成轨迹用于策略学习;L.3 智能体-代理共进化,实时环境证据持续更新代理与智能体,实现双向演进。最终,本文重构世界建模为以智能体为中心的范式,为构建助力智能体更好规划、更快学习、持续进化的世界代理提供了路线图。

原文摘要 · Abstract (English)

Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling offers a natural intermediate proxy that allows agents to query lower-cost, more controllable feedback before committing to real actions. Classical world models instantiate this proxy primarily through future physical-state prediction, a formulation useful yet narrow for agents that require actionable feedback beyond raw state transitions. In this work, we conceptualize Agent-Centric Interactive World Proxies, shifting the fundamental paradigm from physical state transitions to agent-usable information transitions, such as execution outcomes, retrieved experiences or skills, and verification signals, broadening the scope of world modeling to provide versatile feedback for continually improving agents. To systematically map this design space, we organize world proxies into six functional forms based on their feedback modalities: dynamics, spatial, execution, memory/experience, skill, and reward/verification proxies, which together characterize the primary ways world modeling serves agent improvement. We further analyze how these proxies empower agents across three progressive levels: L.1 Inference-Time Guidance, where proxy outputs enrich in-context information for superior decisions; L.2 Training-Time Optimization, where proxy outputs yield rewards, critiques, or synthetic rollouts for policy learning; and L.3 Agent-Proxy Co-Evolution, where real-environment evidence continuously updates both the proxy and the agent for co-evolution. Ultimately, this work recasts world modeling into an agent-centric paradigm, establishing a roadmap for building world proxies that empower agents to plan better, learn faster, and evolve continually.

世界模型智能体持续学习仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。