arXiv:2409.17262cs.RO2024-09被引 10

用跨模态注意力融合视觉与运动数据,让四足机器人自适应复杂地形。

CROSS-GAiT: Cross-Attention-Based Multimodal Representation Fusion for Parametric Gait Adaptation in Complex Terrains

  • 通过视觉和运动数据的跨注意力融合,动态调整步高和髋部张角。
  • 在复杂地形上降低7.04%惯性能量密度,关节总力矩减少27.3%。
  • 适合做机器人自适应行走、多模态感知与控制的研究者参考。

我们提出CROSS-GAiT,一种用于四足机器人的新算法,利用跨注意力机制融合来自视觉和时序输入(线性加速度、角速度、关节力矩)的地形表征。这些融合后的表征被用于持续调节两个关键步态参数(步高和髋部张角),实现对复杂地形的动态响应。视觉输入通过掩码视觉变换器(ViT)编码器处理,时序数据则通过扩张因果卷积编码器处理。跨注意力机制从各模态中选择并整合最相关特征,结合地形特性与机器人动力学,实现智能步态适应。该融合表示使系统能实时应对不可预测地形变化。我们在沥青、混凝土、砖石路、草地、密集植被、碎石、砾石和沙地等多种地形上训练,并在未见过环境中验证其泛化能力。硬件部署于Ghost Robotics Vision 60平台,在高密度植被、不稳定表面、沙丘及可变形地表等挑战场景中表现优异。相比现有方法,惯性测量单元(IMU)能量密度降低至少7.04%,关节总力矩减少27.3%,显著提升稳定性与能效。同时,在四个复杂场景中成功率提升至少64.5%,到达目标时间减少4.91%。此外,学习到的表征在地形分类任务上优于当前最优方法4.48%。

原文摘要 · Abstract (English)

We present CROSS-GAiT, a novel algorithm for quadruped robots that uses Cross Attention to fuse terrain representations derived from visual and time-series inputs; including linear accelerations, angular velocities, and joint efforts. These fused representations are used to continuously adjust two critical gait parameters (step height and hip splay), enabling adaptive gaits that respond dynamically to varying terrain conditions. To generate terrain representations, we process visual inputs through a masked Vision Transformer (ViT) encoder and time-series data through a dilated causal convolutional encoder. The Cross Attention mechanism then selects and integrates the most relevant features from each modality, combining terrain characteristics with robot dynamics for informed gait adaptation. This fused representation allows CROSS-GAiT to continuously adjust gait parameters in response to unpredictable terrain conditions in real-time. We train CROSS-GAiT on a diverse set of terrains including asphalt, concrete, brick pavements, grass, dense vegetation, pebbles, gravel, and sand and validate its generalization ability on unseen environments. Our hardware implementation on the Ghost Robotics Vision 60 demonstrates superior performance in challenging terrains, such as high-density vegetation, unstable surfaces, sandbanks, and deformable substrates. We observe at least a 7.04% reduction in IMU energy density and a 27.3% reduction in total joint effort, which directly correlates with increased stability and reduced energy usage when compared to state-of-the-art methods. Furthermore, CROSS-GAiT demonstrates at least a 64.5% increase in success rate and a 4.91% reduction in time to reach the goal in four complex scenarios. Additionally, the learned representations perform 4.48% better than the state-of-the-art on a terrain classification task.

四足机器人跨模态融合步态适应注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。