arXiv:2608.16164cs.AIcs.RO2026-08

让机器人自动设计训练路径,高效学会在复杂地形上行走。

Trajectory-Level Automatic Curriculum Learning for Legged Locomotion on Unstructured Terrain

论文配图:Trajectory-Level Automatic Curriculum Learning for Legged Locomotion on Unstructured Terrain
图 1 · 摘自论文原文
  • 根据地形地图自动生成训练任务,动态匹配政策进展。
  • 相比直接训练,成功率提升56.3%;比人工设计课程高18.5%以上。
  • 适合研究机器人自主导航、强化学习策略训练的团队使用。

在复杂非结构化地形上训练足式机器人行走策略,需要课程引导以避免早期探索失败。然而,由于非结构化地形缺乏明确的难度排序,现有方法依赖对参数化地形的启发式课程设计,这限制了泛化能力,导致策略过度适应固定的感知模式。为此,我们提出 extbf{ heirname{}}, 一种基于轨迹的自动课程学习框架,直接从非结构化地形地图生成训练任务。每次课程更新时,评估器学习当前策略的难度函数,将给定轨迹任务映射为难度评分;采样器则依据该评分生成新轨迹作为下一阶段策略的课程。这一闭环机制使课程随策略演进持续匹配。定量与定性实验表明, heirname{} 能在非结构化地形上持续提供有效课程,相比无课程的直接训练,轨迹成功率达56.3%的提升;相较手工设计课程,在最困难地形任务上成功率达18.5%提升,且在相同障碍物不同进入方向测试中最高提升达39.74%。

原文摘要 · Abstract (English)

Training locomotion policies for complex unstructured terrain requires a curriculum to avoid early exploration failures. However, since unstructured terrain lacks explicit difficulty ordering for curriculum design, existing methods resort to heuristic curricula over parameterized terrains. This abstraction limits generalization, as policies can overadapt to near-fixed perceptual patterns. To address this, we propose \textbf{\ourname{}}, an \textbf{T}rajectory-level \textbf{A}utomatic \textbf{C}urriculum \textbf{L}earning framework that generates training tasks directly from unstructured terrain maps. At each curriculum update, the evaluator learns a difficulty function for the current policy that maps a given trajectory task to a difficulty score. The sampler then proposes new trajectories guided by the learned evaluator as the curriculum for the next policy update. This forms a closed loop in which the curriculum is iteratively matched to the evolving policy. Quantitative and qualitative experiments show that \ourname{} continuously provides effective curricula on unstructured terrain, improving trajectory success rate by \(56.3\%\) over direct training without curriculum. Compared with handcrafted curriculum learning, our method improves success rate by \(18.5\%\) on the hardest terrain tasks and by up to \(39.74\%\) when evaluating traversal from diverse approach directions on the same obstacle type.

机器人控制强化学习自动课程足式行走

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。