分层策略让四足机器人更智能地跨越复杂地形
Task-Level Decisions to Gait Level Control: A Hierarchical Policy Approach for Quadruped Navigation
- 高层决策根据环境线索选择任务,低层生成对应步态
- 在混合地形和分布外测试中成功率显著提升
- 支持运行时调参与故障诊断,适合实际部署
现实世界中的四足机器人导航受限于高层决策与底层步态执行之间的尺度不匹配,以及在分布外环境变化下的不稳定性。本文提出一种分层策略架构——任务级决策到步态控制(TDGC),通过强化学习在仿真中训练的底层策略,将任务需求映射为可控的行为参数,实现鲁棒的步态生成与平滑切换;高层策略基于稀疏的语义或几何地形线索做出任务导向决策,并转化为底层目标,构建可追溯的决策流程,无需密集地图或高分辨率地形重建。与端到端方法不同,该架构提供明确的运行时接口,支持部署时调参、故障诊断与策略优化。我们引入基于性能的结构化课程,逐步扩展环境难度与扰动范围。实验表明,在混合地形及分布外测试中,任务成功率达到更高水平。
原文摘要 · Abstract (English)
Real-world quadruped navigation is constrained by a scale mismatch between high-level navigation decisions and low-level gait execution, as well as by instabilities under out-of-distribution environmental changes. Such variations challenge sim-to-real transfer and can trigger falls when policies lack explicit interfaces for adaptation. In this paper, we present a hierarchical policy architecture for quadrupedal navigation, termed Task-level Decision to Gait Control (TDGC). A low-level policy, trained with reinforcement learning in simulation, delivers gait-conditioned locomotion and maps task requirements to a compact set of controllable behavior parameters, enabling robust mode generation and smooth switching. A high-level policy makes task-centric decisions from sparse semantic or geometric terrain cues and translates them into low-level targets, forming a traceable decision pipeline without dense maps or high-resolution terrain reconstruction. Different from end-to-end approaches, our architecture provides explicit interfaces for deployment-time tuning, fault diagnosis, and policy refinement. We introduce a structured curriculum with performance-driven progression that expands environmental difficulty and disturbance ranges. Experiments show higher task success rates on mixed terrains and out-of-distribution tests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。