arXiv:2601.17440cs.RO2026-01被引 3

让机器人在复杂环境中同时稳走会抓,靠感知与统一控制

PILOT: A Perceptive Integrated Low-level Controller for Loco-manipulation over Unstructured Scenes

  • 用跨模态编码融合身体感知和视觉信息,精准定位脚步
  • 单阶段强化学习框架实现行走与操作的协同控制
  • 适合需要在非结构化环境作业的人形机器人研究者

人形机器人在人类环境中的多样化交互和服务任务具有巨大潜力,亟需将精确行走与灵巧操作无缝整合的控制器。然而,现有全身控制器普遍缺乏对外部环境的感知能力,难以在复杂非结构化场景中稳定执行任务。为此,我们提出PILOT,一种面向感知式行走-操作一体化的统一单阶段强化学习框架,将感知行走与全身运动控制整合于单一策略中。为增强地形感知并确保精准足位放置,我们设计了跨模态上下文编码器,融合基于预测的本体感觉特征与基于注意力的感知表征。此外,引入专家混合(Mixture-of-Experts)策略架构,以协调多种运动技能,提升不同运动模式下的专业化表现。在仿真及物理实体Unitree G1人形机器人上的大量实验验证了该框架的有效性。相较于现有基线,PILOT展现出更优的稳定性、指令跟踪精度与地形通过能力,凸显其作为非结构化场景下行走-操作底层控制器的鲁棒性潜力。

原文摘要 · Abstract (English)

Humanoid robots hold great potential for diverse interactions and daily service tasks within human-centered environments, necessitating controllers that seamlessly integrate precise locomotion with dexterous manipulation. However, most existing whole-body controllers lack exteroceptive awareness of the surrounding environment, rendering them insufficient for stable task execution in complex, unstructured scenarios.To address this challenge, we propose PILOT, a unified single-stage reinforcement learning (RL) framework tailored for perceptive loco-manipulation, which synergizes perceptive locomotion and expansive whole-body control within a single policy. To enhance terrain awareness and ensure precise foot placement, we design a cross-modal context encoder that fuses prediction-based proprioceptive features with attention-based perceptive representations. Furthermore, we introduce a Mixture-of-Experts (MoE) policy architecture to coordinate diverse motor skills, facilitating better specialization across distinct motion patterns. Extensive experiments in both simulation and on the physical Unitree G1 humanoid robot validate the efficacy of our framework. PILOT demonstrates superior stability, command tracking precision, and terrain traversability compared to existing baselines. These results highlight its potential to serve as a robust, foundational low-level controller for loco-manipulation in unstructured scenes.

人形机器人强化学习行走操作感知控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。