多视角深度感知提升四足机器人动态行走的鲁棒性
Beyond Egocentric Limits: Multi-View Depth-Based Learning for Robust Quadrupedal Locomotion
- 融合自身感知与双路深度视觉,实现多视角环境理解
- 在越障、下台阶等动作中性能优于单视角基线
- 适合需要高鲁棒性的复杂地形机器人应用
近年来,腿式机器人的运动能力已达到类似生物体的动态与特技级表现。然而,现有方法主要依赖机器人自身的第一视角感知,当视角被遮挡时性能受限。本文提出一种基于多视角深度信息的运动控制框架,结合自身(egocentric)与外部(exocentric)观测,增强复杂运动中的环境感知。通过教师-学生知识蒸馏,学生策略学习融合本体感觉与双路深度输入,对真实传感器噪声具有鲁棒性。进一步引入广泛领域随机化,包括随机远程摄像头失效和三维位置扰动,模拟空中-地面协同感知。仿真结果表明,多视角策略在越障、踏步下降等动态动作中显著优于单视角基线,且在外部摄像头部分或完全失效时仍保持稳定。额外实验显示,训练中加入适度视角错位可良好容忍实际偏差。研究证明异构视觉反馈能有效提升四足机器人运动的鲁棒性与敏捷性。为保障可复现性,代码已公开于https://anonymous.4open.science/r/multiview-parkour-6FB8。
原文摘要 · Abstract (English)
Recent progress in legged locomotion has allowed highly dynamic and parkour-like behaviors for robots, similar to their biological counterparts. Yet, these methods mostly rely on egocentric (first-person) perception, limiting their performance, especially when the viewpoint of the robot is occluded. A promising solution would be to enhance the robot's environmental awareness by using complementary viewpoints, such as multiple actors exchanging perceptual information. Inspired by this idea, this work proposes a multi-view depth-based locomotion framework that combines egocentric and exocentric observations to provide richer environmental context during agile locomotion. Using a teacher-student distillation approach, the student policy learns to fuse proprioception with dual depth streams while remaining robust to real-world sensing imperfections. To further improve robustness, we introduce extensive domain randomization, including stochastic remote-camera dropouts and 3D positional perturbations that emulate aerial-ground cooperative sensing. Simulation results show that multi-viewpoints policies outperform single-viewpoint baseline in gap crossing, step descent, and other dynamic maneuvers, while maintaining stability when the exocentric camera is partially or completely unavailable. Additional experiments show that moderate viewpoint misalignment is well tolerated when incorporated during training. This study demonstrates that heterogeneous visual feedback improves robustness and agility in quadrupedal locomotion. Furthermore, to support reproducibility, the implementation accompanying this work is publicly available at https://anonymous.4open.science/r/multiview-parkour-6FB8
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。