arXiv:2512.00005cs.RO2025-12

用想象中的世界模型让机器人少试错,高效探索未知环境。

DREAMer-VXS: A Latent World Model for Sample-Efficient AGV Exploration in Stochastic, Unobserved Environments

  • 用潜空间模型从部分激光雷达数据中学习环境结构和动态
  • 只需真实交互10%次数就达到专家级表现,探索效率提升45%
  • 适合需要少样本、强鲁棒性的自动驾驶或机器人探索任务

基于学习的机器人技术前景广阔,但传统无模型强化学习算法存在样本效率低、易失效的问题。本文提出DREAMer-VXS,一种面向自主地面车辆(AGV)探索的模型基础框架,通过从想象的潜空间轨迹中规划实现高效学习。该方法基于部分且高维的激光雷达观测,构建包含卷积变分自编码器(VAE)与循环状态空间模型(RSSM)的综合世界模型,分别捕捉环境结构与复杂时序动态。利用此模型作为高速模拟器,智能体可几乎完全在想象中训练导航策略,从而将达到专家级性能所需的真实环境交互减少90%,显著优于当前最优的无模型SAC基线。策略由行为-评价网络优化,采用融合任务目标与内在好奇心奖励的复合函数,推动系统性探索未知区域。大量仿真实验表明,DREAMer-VXS不仅学习速度提升数个数量级,且策略更具泛化性和鲁棒性,在未见环境中探索效率提高45%,对动态障碍物也表现出更强适应性。

原文摘要 · Abstract (English)

The paradigm of learning-based robotics holds immense promise, yet its translation to real-world applications is critically hindered by the sample inefficiency and brittleness of conventional model-free reinforcement learning algorithms. In this work, we address these challenges by introducing DREAMer-VXS, a model-based framework for Autonomous Ground Vehicle (AGV) exploration that learns to plan from imagined latent trajectories. Our approach centers on learning a comprehensive world model from partial and high-dimensional LiDAR observations. This world model is composed of a Convolutional Variational Autoencoder (VAE), which learns a compact representation of the environment's structure, and a Recurrent State-Space Model (RSSM), which models complex temporal dynamics. By leveraging this learned model as a high-speed simulator, the agent can train its navigation policy almost entirely in imagination. This methodology decouples policy learning from real-world interaction, culminating in a 90% reduction in required environmental interactions to achieve expert-level performance when compared to state-of-the-art model-free SAC baselines. The agent's behavior is guided by an actor-critic policy optimized with a composite reward function that balances task objectives with an intrinsic curiosity bonus, promoting systematic exploration of unknown spaces. We demonstrate through extensive simulated experiments that DREAMer-VXS not only learns orders of magnitude faster but also develops more generalizable and robust policies, achieving a 45% increase in exploration efficiency in unseen environments and superior resilience to dynamic obstacles.

机器人探索世界模型样本效率强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。