用双系统架构让机器人更高效执行复杂任务
A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
- 分两层处理:大模型慢决策,小模型快执行
- 在RoboCasa上推理更快,成功率更高
- 适合需要实时响应的智能机器人场景
视觉-语言-动作(VLA)模型因其能融合视觉信息与语言指令,使机器人完成复杂任务而受到关注。但现有模型计算量大,难以实现实时性能。为此,我们提出双过程VLA(DP-VLA),受双过程理论启发,采用层级框架:大型系统2模型(L-Sys2)负责复杂推理与决策,小型系统1模型(S-Sys1)处理实时运动控制与感知。借助视觉-语言模型(VLMs),L-Sys2以低频运行,降低计算开销;S-Sys1则确保快速精准的任务执行。在RoboCasa数据集上的实验表明,DP-VLA实现更快推理速度和更高任务成功率,为高级机器人应用提供可扩展解决方案。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models are receiving increasing attention for their ability to enable robots to perform complex tasks by integrating visual context with linguistic commands. However, achieving efficient real-time performance remains challenging due to the high computational demands of existing models. To overcome this, we propose Dual Process VLA (DP-VLA), a hierarchical framework inspired by dual-process theory. DP-VLA utilizes a Large System 2 Model (L-Sys2) for complex reasoning and decision-making, while a Small System 1 Model (S-Sys1) handles real-time motor control and sensory processing. By leveraging Vision-Language Models (VLMs), the L-Sys2 operates at low frequencies, reducing computational overhead, while the S-Sys1 ensures fast and accurate task execution. Experimental results on the RoboCasa dataset demonstrate that DP-VLA achieves faster inference and higher task success rates, providing a scalable solution for advanced robotic applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。