统一视觉语言动作模型,让机器人长时任务更智能
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
- 用统一框架同时生成任务分解和机器人动作
- 在Ravens模拟器上达成90%以上任务成功率
- 适合研究具身智能与多步决策的学者
现实世界中的具身智能体需应对长时程任务,这类任务要求高阶目标分解与多步动作规划。现有视觉语言动作(VLA)模型在规划能力上不足,而分层架构存在协调问题。本文提出统一的LoHoVLA框架,以预训练视觉语言模型为骨干,联合生成语言子任务与动作指令,共享表征提升泛化性;同时采用分层闭环控制机制,减少规划与执行误差。为训练该模型,构建了基于Ravens模拟器的LoHoSet数据集,包含20个长时程任务,每任务1,000条专家演示,涵盖视觉观测、语言目标、子任务与机器人动作。实验表明,LoHoVLA在Ravens模拟器上的长时程任务中显著优于传统分层与标准VLA方法。
原文摘要 · Abstract (English)
Real-world embodied agents face long-horizon tasks, characterized by high-level goals demanding multi-step solutions beyond single actions. Successfully navigating these requires both high-level task planning (i.e., decomposing goals into sub-tasks) and low-level motion control (i.e., generating precise robot actions). While existing vision language action (VLA) models and hierarchical architectures offer potential in embodied tasks, the former often falter in planning, and the latter can suffer from coordination issues, both hampering performance. We introduce a new unified VLA framework for long-horizon tasks, dubbed LoHoVLA, to overcome these limitations. LoHoVLA leverages a large pretrained vision language model (VLM) as the backbone to jointly generate language and action tokens for sub-task generation and robot action prediction, respectively. This shared representation promotes better generalization across tasks. Additionally, LoHoVLA embraces a hierarchical closed-loop control mechanism to mitigate errors originating from both high-level planning and low-level control. To train LoHoVLA, we introduce LoHoSet, a dataset built on the Ravens simulator, containing 20 long-horizon tasks, each with 1,000 expert demonstrations composed of visual observations, linguistic goals, sub-tasks, and robot actions. Experimental results show that LoHoVLA significantly surpasses both hierarchical and standard VLA approaches on long-horizon embodied tasks in the Ravens simulator. These findings underscore the promise of unified architectures for advancing generalizable embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。