直接从深度图生成行走控制,无需中间几何抽象。
CReF: Cross-modal and Recurrent Fusion for Depth-conditioned Humanoid Locomotion
- 用跨模态注意力融合本体感知与深度信息,端到端学习。
- 在多种复杂地形上实现零样本迁移,包括反光和杂乱环境。
- 适合希望跳过地图构建、直接从深度图驱动机器人的研究者。
在复杂地形上稳定行走日益依赖外部感知,但现有方法常依赖显式的几何抽象,如通过机器人中心的2.5D地形表示或辅助几何目标来引导深度学习。这些方法引入额外的地图构建或多阶段技能迁移流程。本文提出CReF(Cross-modal and Recurrent Fusion),一种单阶段深度条件化人形机器人行走框架,直接从前向深度图学习行走相关特征,无需显式几何中间表示。CReF通过本体感知查询的跨模态注意力耦合本体感知与深度标记,使用门控残差融合模块融合表征,并利用由高速公路式输出门调控的门控循环单元(GRU)进行时序整合,实现状态依赖的递归与前馈特征融合。为增强地形交互,引入地形感知足点放置奖励,从脚部点云样本中提取可支撑足点候选,并奖励落在最近可支撑候选附近的触地位置。仿真与物理人形机器人实验表明,CReF可在多样地形上实现稳健穿越,并有效实现零样本迁移至包含扶手、空心托盘、严重反射干扰及视觉杂乱的室外场景。
原文摘要 · Abstract (English)
Stable traversal over geometrically complex terrain increasingly requires exteroceptive perception, yet prior perceptive humanoid locomotion methods often remain tied to explicit geometric abstractions, either by mediating control through robot-centric 2.5D terrain representations or by shaping depth learning with auxiliary geometry-related targets. While effective, these approaches introduce additional map-construction procedures or multi-stage skill-transfer processes beyond direct depth-to-control learning. We propose CReF (Cross-modal and Recurrent Fusion), a single-stage depth-conditioned humanoid locomotion framework that learns locomotion-relevant features directly from raw forward-facing depth without explicit geometric intermediates. CReF couples proprioception and depth tokens through proprioception-queried cross-modal attention, fuses the resulting representation with a gated residual fusion block, and performs temporal integration with a Gated Recurrent Unit (GRU) regulated by a highway-style output gate for state-dependent blending of recurrent and feedforward features. To further improve terrain interaction, we introduce a terrain-aware foothold placement reward that extracts supportable foothold candidates from foot-end point-cloud samples and rewards touchdown locations that lie close to the nearest supportable candidate. Experiments in simulation and on a physical humanoid demonstrate robust traversal over diverse terrains and effective zero-shot transfer to real-world scenes containing handrails, hollow pallet assemblies, severe reflective interference, and visually cluttered outdoor surroundings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。