用视觉语言模型+实时地形感知,自动设计奖励函数让机器人爬楼梯
E-SDS: Environment-aware See it, Do it, Sorted - Automated Environment-Aware Reinforcement Learning for Humanoid Locomotion
- 结合视觉语言模型与实时地形分析,自动生成环境感知奖励函数
- 在四类地形上使速度追踪误差降低51.9%~82.6%,首次实现稳定下楼梯
- 将人工调奖时间从数天缩短至两小时以内,适合具身智能研究者
视觉语言模型(VLMs)在自动化人形机器人运动中的奖励设计方面展现出潜力,可避免繁琐的人工工程。然而,现有基于VLM的方法本质上是“盲的”,缺乏导航复杂地形所需的环境感知能力。本文提出E-SDS(Environment-aware See it, Do it, Sorted),通过将VLM与实时地形传感器分析相结合,自动生成基于示例视频的奖励函数,从而训练出具备环境感知能力的稳健运动策略。在Unitree G1人形机器人上评估的四种地形(简单、间隙、障碍物、台阶)中,E-SDS唯一实现了成功下楼梯,而采用手动设计奖励或非感知型自动化基线的策略均无法完成该任务。所有地形下,E-SDS的速度追踪误差降低了51.9%~82.6%。该框架将人工奖励设计时间从数天减少至不足两小时,同时生成更鲁棒、更强大的运动策略。
原文摘要 · Abstract (English)
Vision-language models (VLMs) show promise in automating reward design in humanoid locomotion, which could eliminate the need for tedious manual engineering. However, current VLM-based methods are essentially "blind", as they lack the environmental perception required to navigate complex terrain. We present E-SDS (Environment-aware See it, Do it, Sorted), a framework that closes this perception gap. E-SDS integrates VLMs with real-time terrain sensor analysis to automatically generate reward functions that facilitate training of robust perceptive locomotion policies, grounded by example videos. Evaluated on a Unitree G1 humanoid across four distinct terrains (simple, gaps, obstacles, stairs), E-SDS uniquely enabled successful stair descent, while policies trained with manually-designed rewards or a non-perceptive automated baseline were unable to complete the task. In all terrains, E-SDS also reduced velocity tracking error by 51.9-82.6%. Our framework reduces the human effort of reward design from days to less than two hours while simultaneously producing more robust and capable locomotion policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。