arXiv:2506.22827cs.RO2025-06中稿 · the RSS 2025 Works…被引 8

用视觉语言模型实现人形机器人多步骤操作的分层规划

Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation

  • 分三层:底层控制、中层技能策略、高层视觉语言规划
  • 真实世界40次试验成功率达73%,可实时监控任务进展
  • 适合需要复杂操作的工业与家庭服务机器人研究者

使人形机器人可靠执行复杂的多步骤操作任务,对其实现在工业和家庭环境中的有效部署至关重要。本文提出一种分层规划与控制框架,实现可靠的多步骤人形机器人操作。该系统包含三层:(1) 基于强化学习的底层控制器,负责跟踪全身运动目标;(2) 通过模仿学习训练的中层技能策略,生成任务各步骤的运动目标;(3) 高层视觉语言规划模块,利用预训练的视觉语言模型(VLMs)决定应执行哪些技能,并实时监控其完成情况。在Unitree G1人形机器人上对非抓取式拾放任务进行了实验验证。在40次真实世界测试中,该分层系统成功完成了完整操作序列,成功率高达73%。实验验证了所提框架的可行性,突显了基于VLM的技能规划与监控在多步骤操作场景中的优势。视频演示见 https://vlp-humanoid.github.io/。

原文摘要 · Abstract (English)

Enabling humanoid robots to reliably execute complex multi-step manipulation tasks is crucial for their effective deployment in industrial and household environments. This paper presents a hierarchical planning and control framework designed to achieve reliable multi-step humanoid manipulation. The proposed system comprises three layers: (1) a low-level RL-based controller responsible for tracking whole-body motion targets; (2) a mid-level set of skill policies trained via imitation learning that produce motion targets for different steps of a task; and (3) a high-level vision-language planning module that determines which skills should be executed and also monitors their completion in real-time using pretrained vision-language models (VLMs). Experimental validation is performed on a Unitree G1 humanoid robot executing a non-prehensile pick-and-place task. Over 40 real-world trials, the hierarchical system achieved a 73% success rate in completing the full manipulation sequence. These experiments confirm the feasibility of the proposed hierarchical system, highlighting the benefits of VLM-based skill planning and monitoring for multi-step manipulation scenarios. See https://vlp-humanoid.github.io/ for video demonstrations of the policy rollout.

人形机器人视觉语言多步操作分层规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。