通过梯度优化生成渐进式环境,提升机器人导航的泛化能力
Curriculum Generation under Structured Parametric Environments for Robust Navigation Policies

- 用梯度法自动设计逐步变难的训练环境
- 在两个连续控制任务中均显著优于基线方法
- 适合需要强鲁棒性的自主导航系统研究
自主智能体的稳健导航策略需适应不断变化的环境条件,如转向速率、障碍物、摩擦力、坑洞和坡度。课程生成为提升泛化能力提供了一种合理机制,通过逐步调整训练环境实现,但如何高效自动化地设计此类课程仍具挑战。本文提出一种基于单向梯度优化的可重参数化课程生成框架,用于结构化连续环境参数。为增强多模态观测空间(图像与标量输入)中的鲁棒性,引入分布偏移正则化目标,以促进更精细的潜在表征学习。在两个连续控制环境——2D障碍赛车变体和双足行走者变体——上进行评估,其中耦合的环境参数共同影响策略性能。在五个随机种子下,该方法始终优于原始策略训练、随机参数采样、手动课程、前沿引导方法、自定步长强化学习(SPRL)、带高斯混合模型的绝对学习进度(ALP-GMM)以及逆向课程学习基线。消融实验进一步验证了重参数化课程机制的有效性,并揭示了正则化目标在不同环境中的差异化收益。
原文摘要 · Abstract (English)
Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, friction, pits, and slopes. Curriculum generation provides a principled mechanism for improving generalization by progressively adapting training environments, but designing such curricula in a sample-efficient and automated manner remains challenging. This paper proposes a reparameterized curriculum generation framework for structured continuous environment parameters using unidirectional gradient-based optimization. To improve robustness in multimodal observation spaces consisting of image-based and scalar inputs, a distribution-shift regularization objective is incorporated to encourage the learning of finer-grained latent representations. The proposed method is evaluated across two continuous-control OpenAI Gym environments: a 2D obstacle-based Car Racing variant and Bipedal Walker variant, where coupled environment parameters jointly influence policy performance. Across five random seeds, our method consistently outperforms vanilla policy training, random parameter sampling, manual curricula, frontier-based methods, Self-Paced Reinforcement Learning (SPRL), Absolute Learning Progress with Gaussian Mixture Models (ALP-GMM), and reverse curriculum learning baselines. Ablation studies further demonstrate the effectiveness of the reparameterized curriculum mechanism across both environments, while highlighting environment-dependent benefits of the auxiliary regularization objective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。