端到端导航让仿人机器人安全舒适地穿行复杂环境
End-to-End Humanoid Robot Safe and Comfortable Locomotion Policy
- 直接用雷达点云生成运动指令,无需感知中间步骤
- 通过约束强化学习确保避障安全,同时提升动作流畅度
- 适合需要人机共存场景的机器人研发者参考
将仿人机器人部署于非结构化、以人为中心的环境,需具备超越基础行走能力的导航能力,包括鲁棒感知、可证明的安全性以及社会兼容行为。当前强化学习方法常受限于缺乏环境感知的盲控策略,或在复杂三维障碍物前失效的视觉系统。本文提出一种端到端的行走策略,直接将原始时空雷达点云映射为运动指令,实现在杂乱动态场景中的稳健导航。我们将控制问题建模为约束马尔可夫决策过程(CMDP),以形式化分离安全性与任务目标。核心贡献是将控制屏障函数(CBFs)原理转化为CMDP中的惩罚项,使无模型的罚近端策略优化(P3O)在训练中自动遵守安全约束。此外,我们引入基于人机交互研究的舒适性奖励,促进动作平滑、可预测且低侵扰。通过成功实现从仿真到真实仿人机器人的迁移,验证了框架的有效性,机器人可在静态与动态三维障碍物间敏捷、安全地穿行。
原文摘要 · Abstract (English)
The deployment of humanoid robots in unstructured, human-centric environments requires navigation capabilities that extend beyond simple locomotion to include robust perception, provable safety, and socially aware behavior. Current reinforcement learning approaches are often limited by blind controllers that lack environmental awareness or by vision-based systems that fail to perceive complex 3D obstacles. In this work, we present an end-to-end locomotion policy that directly maps raw, spatio-temporal LiDAR point clouds to motor commands, enabling robust navigation in cluttered dynamic scenes. We formulate the control problem as a Constrained Markov Decision Process (CMDP) to formally separate safety from task objectives. Our key contribution is a novel methodology that translates the principles of Control Barrier Functions (CBFs) into costs within the CMDP, allowing a model-free Penalized Proximal Policy Optimization (P3O) to enforce safety constraints during training. Furthermore, we introduce a set of comfort-oriented rewards, grounded in human-robot interaction research, to promote motions that are smooth, predictable, and less intrusive. We demonstrate the efficacy of our framework through a successful sim-to-real transfer to a physical humanoid robot, which exhibits agile and safe navigation around both static and dynamic 3D obstacles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。