让智能体像人一样想象未知场景,提升决策能力。
Generative World Explorer
- 用生成模型实现虚拟心理探索,获取想象中的观察数据
- 在长时序探索中保持高质量一致的生成结果
- 可提升现有决策模型(如大语言模型)的规划能力
在具身智能中,基于不完整观测进行规划是一个核心挑战。以往方法多依赖智能体物理探索以更新对世界状态的认知。而人类可通过心理想象未见区域并修正信念,无需实时物理探索即可做出更优决策。为此,我们提出生成式世界探索框架(Genex),使智能体能在大型3D环境(如城市场景)中进行心理探索,生成想象中的观测以更新信念,从而辅助当前决策。为训练该框架,我们构建了合成城市场景数据集Genex-DB。实验表明:(1) Genex可在大规模虚拟世界中实现长时序探索,并生成高质量、一致的观测;(2) 利用生成观测更新的信念可有效提升现有决策模型(如大语言模型代理)的规划表现。
原文摘要 · Abstract (English)
Planning with partial observation is a central challenge in embodied AI. A majority of prior works have tackled this challenge by developing agents that physically explore their environment to update their beliefs about the world state. In contrast, humans can $\textit{imagine}$ unseen parts of the world through a mental exploration and $\textit{revise}$ their beliefs with imagined observations. Such updated beliefs can allow them to make more informed decisions, without necessitating the physical exploration of the world at all times. To achieve this human-like ability, we introduce the $\textit{Generative World Explorer (Genex)}$, an egocentric world exploration framework that allows an agent to mentally explore a large-scale 3D world (e.g., urban scenes) and acquire imagined observations to update its belief. This updated belief will then help the agent to make a more informed decision at the current step. To train $\textit{Genex}$, we create a synthetic urban scene dataset, Genex-DB. Our experimental results demonstrate that (1) $\textit{Genex}$ can generate high-quality and consistent observations during long-horizon exploration of a large virtual physical world and (2) the beliefs updated with the generated observations can inform an existing decision-making model (e.g., an LLM agent) to make better plans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。