arXiv:2410.10766cs.ROcs.AI2024-10CoRL被引 8

用扩散模型动态生成复杂地形,让机器人学会在不平坦地面自主导航。

Adaptive Diffusion Terrain Generator for Autonomous Uneven Terrain Navigation

  • 基于去噪扩散模型,根据当前策略表现自适应生成新地形。
  • 训练出的策略在复杂不平地形上表现优于传统生成与自然环境。
  • 适合研究机器人自主导航、强化学习环境生成的学者使用。

无模型强化学习已成为开发鲁棒机器人控制策略的有效方法,可应对复杂非结构化地形。该方法有效性依赖两大要素:(1) 利用大规模并行物理仿真加速策略训练;(2) 环境生成器需构建足够挑战但可达成的地形以持续提升策略性能。现有环境生成方法多依赖参数化启发式规则,限制了地形多样性与真实性。本文提出自适应扩散地形生成器(ADTG),利用去噪扩散概率模型(Denoising Diffusion Probabilistic Models)动态扩展训练环境,生成更丰富复杂的地形,且能根据当前策略表现自适应调整。ADTG通过初始噪声优化引导生成过程,融合来自已有训练环境的噪声污染地形,并按策略在各环境中的表现加权。通过调节噪声污染程度,可无缝切换于相似地形微调与全新地形拓展之间。实验表明,由ADTG训练的策略优于程序生成与自然环境下的策略,也超越多种主流导航方法。

原文摘要 · Abstract (English)

Model-free reinforcement learning has emerged as a powerful method for developing robust robot control policies capable of navigating through complex and unstructured terrains. The effectiveness of these methods hinges on two essential elements: (1) the use of massively parallel physics simulations to expedite policy training, and (2) an environment generator tasked with crafting sufficiently challenging yet attainable terrains to facilitate continuous policy improvement. Existing methods of environment generation often rely on heuristics constrained by a set of parameters, limiting the diversity and realism. In this work, we introduce the Adaptive Diffusion Terrain Generator (ADTG), a novel method that leverages Denoising Diffusion Probabilistic Models to dynamically expand existing training environments by adding more diverse and complex terrains adaptive to the current policy. ADTG guides the diffusion model's generation process through initial noise optimization, blending noise-corrupted terrains from existing training environments weighted by the policy's performance in each corresponding environment. By manipulating the noise corruption level, ADTG seamlessly transitions between generating similar terrains for policy fine-tuning and novel ones to expand training diversity. Our experiments show that the policy trained by ADTG outperforms both procedural generated and natural environments, along with popular navigation methods.

机器人导航扩散模型强化学习地形生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。