让机器人像人一样慢思考,提升复杂操作能力。
Hume: Introducing System-2 Thinking in Visual-Language-Action Model
- 引入双系统架构,用价值引导的慢思考生成动作候选
- 在仿真和真实机器人上均超越现有最先进模型
- 适合需要精细控制的机器人任务研究者
人类在处理物理世界中的复杂任务时,会先进行缓慢的思考再执行动作。这种思维模式近期显著提升了大语言模型在数字领域解决复杂任务的能力,但在与物理世界交互的机器人基础模型中仍鲜有探索。本文提出 Hume:一种具备价值引导式系统2思考与级联动作去噪的视觉-语言-动作(VLA)双系统模型,旨在赋予视觉-语言-动作模型类人的思考能力以实现灵巧机器人控制。系统2通过在视觉-语言-动作模型主干网络上添加新颖的价值查询头,估算所预测动作的状态-动作价值,并通过重复采样多个动作候选,依据价值选择最优动作。系统1是一个轻量级反应式视觉运动策略,接收系统2选定的动作后,进行级联动作去噪以实现灵巧控制。部署时,系统2以低频执行价值引导思考,系统1异步接收动作并实时生成流畅动作。实验表明,Hume在多个仿真基准和真实机器人部署中均优于现有最先进视觉-语言-动作模型。
原文摘要 · Abstract (English)
Humans practice slow thinking before performing actual actions when handling complex tasks in the physical world. This thinking paradigm, recently, has achieved remarkable advancement in boosting Large Language Models (LLMs) to solve complex tasks in digital domains. However, the potential of slow thinking remains largely unexplored for robotic foundation models interacting with the physical world. In this work, we propose Hume: a dual-system Vision-Language-Action (VLA) model with value-guided System-2 thinking and cascaded action denoising, exploring human-like thinking capabilities of Vision-Language-Action models for dexterous robot control. System 2 of Hume implements value-Guided thinking by extending a Vision-Language-Action Model backbone with a novel value-query head to estimate the state-action value of predicted actions. The value-guided thinking is conducted by repeat sampling multiple action candidates and selecting one according to state-action value. System 1 of Hume is a lightweight reactive visuomotor policy that takes System 2 selected action and performs cascaded action denoising for dexterous robot control. At deployment time, System 2 performs value-guided thinking at a low frequency while System 1 asynchronously receives the System 2 selected action candidate and predicts fluid actions in real time. We show that Hume outperforms the existing state-of-the-art Vision-Language-Action models across multiple simulation benchmark and real-robot deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。