用视觉语言模型预判地面摩擦,让轮足机器人更安全地在滑地行走
Friction-Aware Safety Locomotion for Wheeled-legged Robots using Vision Language Models and Reinforcement Learning
- 结合视觉语言模型与强化学习,提前预测地面摩擦系数
- 在模拟和真实轮式倒立摆上实现滑地轨迹跟踪,基线方法失败
- 适合做复杂地形移动的机器人研发,尤其关注安全控制的团队
控制轮足机器人在滑地上运动极具挑战性,因其依赖持续的地面接触。与四足或双足机器人不同,轮足机器人在出现瞬时打滑时极易失稳且难以恢复。若能在接触前预判地面物理特性(如摩擦系数),则可主动调整控制策略以降低打滑风险。本文提出一种摩擦感知的安全步态框架,融合视觉语言模型(VLM)与强化学习(RL)策略。采用检索增强生成(RAG)方法估计摩擦系数(CoF),并将其显式纳入RL策略中,使机器人可在接触前根据预测的摩擦条件自适应调整速度。该框架在仿真环境和一台定制的轮式倒立摆(WIP)平台上验证。实验结果表明,本方法成功完成滑地上的轨迹跟踪任务,而仅依赖本体感知反馈的基线方法均失败。这凸显了显式预测并利用地面摩擦信息对安全步态的重要性,也指明了利用VLM评估地面状况这一具有前景的研究方向,这对纯视觉方法仍具挑战。
原文摘要 · Abstract (English)
Controlling Wheeled-legged robots is challenging especially on slippery surfaces due to their dependence on continuous ground contact. Unlike quadrupeds or bipeds, which can leverage multiple fixed contact points for recovery, wheeled-legged robots are highly susceptible to slip, where even momentary loss of traction can result in irrecoverable instability. Anticipating ground physical properties such as friction before contact would allow proactive control adjustments, reducing slip risk. In this paper, we propose a friction-aware safety locomotion framework that integrates Vision-Language Models (VLMs) with a Reinforcement Learning (RL) policy. Our method employs a Retrieval-Augmented Generation (RAG) approach to estimate the Coefficient of Friction (CoF), which is then explicitly incorporated into the RL policy. This enables the robot to adapt its speed based on predicted friction conditions before contact. The framework is validated through experiments in both simulation and on a physical customized Wheeled Inverted Pendulum (WIP). Experimental results show that our approach successfully completes trajectory tracking tasks on slippery surfaces, whereas baseline methods relying solely on proprioceptive feedback fail. These findings highlight the importance and effectiveness of explicitly predicting and utilizing ground friction information for safe locomotion. They also point to a promising research direction in exploring the use of VLMs for estimating ground conditions, which remains a significant challenge for purely vision-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。