arXiv:2606.00637cs.RO2026-06被引 1

提出一种分解全局与局部注意力的地形编码方法,提升人形机器人在复杂地形上的感知行走能力。

Global-Local Attention Decomposition for Terrain Encoding in Humanoid Perceptive Locomotion

论文配图:Global-Local Attention Decomposition for Terrain Encoding in Humanoid Perceptive Locomotion
图 1 · 摘自论文原文
  • 将地形感知分为全局上下文与局部关键点两支路,分别用注意力池化和地形显著性稀疏编码
  • 在跳跃石、台阶等复杂地形上实现稳定行走,且无需额外导航规划
  • 可在真实机器人上零样本迁移,支持基于激光雷达的自主避障与路径跟随

尽管强化学习显著提升了人形机器人的行走能力,但在稀疏落脚点地形和受限环境中仍表现不佳。成功应对这些场景需要兼具广阔的地形感知与精确的落脚点选择,而传统编码器常将二者混淆。为此,我们提出全局-局部注意力分解(GLAD)用于人形机器人地形编码。该方法通过以机器人为中心的高程图实现从粗到细的编码:全局注意力分支使用注意力池化总结周围地形上下文;局部注意力分支则基于地形显著性稀疏化局部特征,并结合状态条件注意力编码与落脚点相关的几何信息。这种显式分解避免了细粒度空间线索的模糊化,同时降低训练开销。实验表明,GLAD可实现对挑战性缝隙、踏板和台阶的可靠行走。此外,学习到的策略展现出涌现的地形响应行为,在仅接收前进速度指令的情况下,能自主沿狭窄路径行走并避开障碍物,无需显式导航规划。在搭载机载激光雷达的真实单位树G1机器人上部署,该方法在多样化的稀疏落脚点与障碍密集场景中实现了鲁棒的零样本模拟到现实迁移。

原文摘要 · Abstract (English)

Although reinforcement learning has significantly advanced humanoid locomotion, perceptive policies still struggle on sparse-foothold terrain and constrained environments. Success in these scenarios requires both broad terrain awareness and precise foothold selection, two perceptual roles that conventional encoders often entangle. To address this challenge, we propose Global-Local Attention Decomposition (GLAD) for terrain encoding in humanoid locomotion. Realized by a coarse-to-fine encoder over a robot-centric elevation map, GLAD explicitly separates these objectives: a global attention branch uses attention pooling to summarize the surrounding terrain context, while a local attention branch sparsifies the local features by terrain saliency and applies state-conditioned attention to encode precise foothold-relevant geometry. This explicit attention decomposition prevents the dilution of fine-grained spatial cues while reducing training overhead. Experiments demonstrate that GLAD enables reliable locomotion over challenging gaps, stepping stones, and stairs. Furthermore, the learned policy exhibits emergent terrain-responsive behaviors, autonomously following narrow paths and avoiding obstacles under forward-velocity commands alone, without explicit navigation planners. In real-world deployment on a Unitree G1 humanoid robot using onboard LiDAR, the proposed method achieves robust zero-shot sim-to-real transfer across diverse sparse-foothold and obstacle-rich domains.

人形机器人地形感知注意力机制强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。