arXiv:2510.16123cs.LG2025-10NeurIPS

不训练直接搜索记忆,实现零样本世界模型

Zero-shot World Models via Search in Memory

  • 用相似性检索和随机表示构建世界模型
  • 长时序预测性能优于传统训练模型
  • 适合快速部署于多类视觉环境

世界模型已广泛应用于强化学习领域,其建模环境动态转移的能力显著提升了在线强化学习的样本效率。以Dreamer为代表的方法在多种图像环境中的表现尤为突出。本文提出一种无需训练的搜索型世界模型,通过记忆中的相似性检索与随机表示来逼近环境动态。我们将其与经典的世界模型PlaNet进行对比,评估了潜在空间重构质量及重建图像的感知相似度,涵盖单步与长时序预测任务。实验结果表明,该搜索型模型在两类任务中均达到与训练型模型相当的性能,尤其在多样视觉环境下长时序预测表现更优。

原文摘要 · Abstract (English)

World Models have vastly permeated the field of Reinforcement Learning. Their ability to model the transition dynamics of an environment have greatly improved sample efficiency in online RL. Among them, the most notorious example is Dreamer, a model that learns to act in a diverse set of image-based environments. In this paper, we leverage similarity search and stochastic representations to approximate a world model without a training procedure. We establish a comparison with PlaNet, a well-established world model of the Dreamer family. We evaluate the models on the quality of latent reconstruction and on the perceived similarity of the reconstructed image, on both next-step and long horizon dynamics prediction. The results of our study demonstrate that a search-based world model is comparable to a training based one in both cases. Notably, our model show stronger performance in long-horizon prediction with respect to the baseline on a range of visually different environments.

世界模型零样本搜索强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。