系统研究分层视觉语言动作模型设计,提升机器人复杂操作能力。
What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents

- 将分层VLA系统统一到选项控制框架下,对比不同设计组合。
- 在仿真和真实ALOHA机器人上,新系统成功率显著高于扁平或简单分层方案。
- 为构建更鲁棒、有原则的分层智能体提供可复用的设计指南。
分层视觉-语言-动作(Hi-VLA)系统作为复杂机器人操作的新兴范式,通过高层视觉语言模型(VLM)规划器将任务分解为语言子目标,由低层VLA控制器执行。尽管已有实证进展,但缺乏统一的设计原则:现有系统在规划器与控制器的选择与连接、切换机制及观测与记忆表示方式上差异显著。本文对Hi-VLA系统设计进行系统性研究,将代表性系统统一至选项风格控制框架,在短时序、长时序及推理密集型任务上基准测试核心设计选择。分析提炼出实用的设计原则,揭示模型选择与接口机制如何共同影响性能。应用这些原则所构建的系统,在仿真与真实ALOHA机器人上均显著优于传统扁平VLA控制或粗略设计的分层结构。结果为构建更强大、鲁棒且有原则的分层VLA智能体奠定基础。更多信息与视频见 jiahenghu.github.io/hi-vla。
原文摘要 · Abstract (English)
Hierarchical vision-language-action (Hi-VLA) systems have emerged as a promising paradigm for complex robot manipulation, by using high-level VLM planners to decompose tasks into language subgoals executed by low-level VLA controllers. Despite recent empirical progress, there is a lack of unified design principles for these systems: existing Hi-VLA systems differ in how they choose and connect planners, controllers, mechanisms to switch between the two, and how observations and memory are represented in the planner. In this paper, we present a systematic study of Hi-VLA design for robot manipulation. We unify representative Hi-VLA agents under an options-style control framework and benchmark core design choices across short-horizon, long-horizon, and reasoning-intensive tasks. Our analysis distills practical principles for building Hi-VLA systems, showing how model choices and interface mechanisms jointly shape performance. Applying these principles yields a substantially stronger system than either flat VLA control or a naively designed hierarchy, across experiments both in simulation and on a real ALOHA robot. Overall, our results provide a foundation for building more capable, robust, and principled hierarchical VLA agents. More information and video at jiahenghu.github.io/hi-vla.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。