arXiv:2506.01759cs.ROcs.SY2025-06被引 2

用扩散模型动态生成复杂环境,提升机器人仿真到现实的迁移效果。

ADEPT: Adaptive Diffusion Environment for Policy Transfer Sim-to-Real

  • 基于去噪扩散模型,根据策略表现自适应生成新环境。
  • 在野外导航任务中,性能超越程序生成与自然环境基准。
  • 适合需要高泛化能力的机器人控制研究者使用。

无模型强化学习已成为开发鲁棒机器人控制策略的强大方法,能够应对复杂非结构化环境。其有效性依赖于两个关键因素:(1) 利用大规模并行物理仿真加速策略训练;(2) 环境生成器需构造足够挑战但可实现的环境以持续提升策略性能。现有室外环境生成方法多依赖参数化启发式规则,限制了多样性与真实性。本文提出 ADEPT——一种零样本仿真到现实迁移的自适应扩散环境生成框架,利用去噪扩散概率模型(Denoising Diffusion Probabilistic Models)动态扩展训练环境,加入更多样、更复杂的环境以适应当前策略。ADEPT通过初始噪声优化引导生成过程,融合来自已有训练环境的噪声污染环境,并按策略在各环境中的表现加权。通过调节噪声污染程度,可无缝切换至相似环境用于策略微调或生成新颖环境以拓展训练多样性。为在非铺装路面导航中评估 ADEPT,我们提出了快速有效的多层地图表示用于野生环境生成。实验表明,由 ADEPT 训练的策略优于程序生成与自然环境基线,以及多种主流导航方法。

原文摘要 · Abstract (English)

Model-free reinforcement learning has emerged as a powerful method for developing robust robot control policies capable of navigating through complex and unstructured environments. The effectiveness of these methods hinges on two essential elements: (1) the use of massively parallel physics simulations to expedite policy training, and (2) an environment generator tasked with crafting sufficiently challenging yet attainable environments to facilitate continuous policy improvement. Existing methods of outdoor environment generation often rely on heuristics constrained by a set of parameters, limiting the diversity and realism. In this work, we introduce ADEPT, a novel \textbf{A}daptive \textbf{D}iffusion \textbf{E}nvironment for \textbf{P}olicy \textbf{T}ransfer in the zero-shot sim-to-real fashion that leverages Denoising Diffusion Probabilistic Models to dynamically expand existing training environments by adding more diverse and complex environments adaptive to the current policy. ADEPT guides the diffusion model's generation process through initial noise optimization, blending noise-corrupted environments from existing training environments weighted by the policy's performance in each corresponding environment. By manipulating the noise corruption level, ADEPT seamlessly transitions between generating similar environments for policy fine-tuning and novel ones to expand training diversity. To benchmark ADEPT in off-road navigation, we propose a fast and effective multi-layer map representation for wild environment generation. Our experiments show that the policy trained by ADEPT outperforms both procedural generated and natural environments, along with popular navigation methods.

强化学习扩散模型仿真到现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。