融合身体感知与视觉信息,让四足机器人更稳地穿越复杂户外环境
KiVi: Kinesthetic-Visuospatial Integration for Dynamic and Safe Egocentric Legged Locomotion
- 分路处理本体感觉与视觉信息,用身体感知保证稳定
- 在遮挡和光照变化下仍能稳定行走,通过记忆注意力增强鲁棒性
- 适合真实场景中对安全与动态适应要求高的四足机器人应用
基于视觉的行走控制在使四足机器人感知并适应复杂环境方面展现出巨大潜力。然而,视觉信息本身易受遮挡、反射和光照变化影响,常导致运动不稳。受动物感觉运动整合启发,我们提出KiVi框架,将本体感觉(kinesthetics)与视觉空间推理(visuospatial reasoning)分离处理:前者编码身体运动的本体感知,后者捕捉周围地形的视觉信息。具体而言,系统以本体感知为稳定主干,选择性引入视觉信息实现地形感知与避障。这种模态平衡但可整合的设计,结合记忆增强注意力机制,使机器人能稳健解读视觉线索,同时通过本体感知实现故障容错。大量实验表明,该方法使四足机器人可在多样地形上稳定行进,在非结构化室外环境中可靠运行,对训练中未见的分布外(OOD)视觉噪声和遮挡仍具鲁棒性,充分体现了其在真实世界腿式行走中的有效性与适用性。
原文摘要 · Abstract (English)
Vision-based locomotion has shown great promise in enabling legged robots to perceive and adapt to complex environments. However, visual information is inherently fragile, being vulnerable to occlusions, reflections, and lighting changes, which often cause instability in locomotion. Inspired by animal sensorimotor integration, we propose KiVi, a Kinesthetic-Visuospatial integration framework, where kinesthetics encodes proprioceptive sensing of body motion and visuospatial reasoning captures visual perception of surrounding terrain. Specifically, KiVi separates these pathways, leveraging proprioception as a stable backbone while selectively incorporating vision for terrain awareness and obstacle avoidance. This modality-balanced, yet integrative design, combined with memory-enhanced attention, allows the robot to robustly interpret visual cues while maintaining fallback stability through proprioception. Extensive experiments show that our method enables quadruped robots to stably traverse diverse terrains and operate reliably in unstructured outdoor environments, remaining robust to out-of-distribution(OOD) visual noise and occlusion unseen during training, thereby highlighting its effectiveness and applicability to real-world legged locomotion. Project Page: https://marmotlab.github.io/kivi-quadruped/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。