分层设计让机器人既会规划又会执行,处理复杂任务更精准。
HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System

- 高阶用视觉语言模型分解任务并定位目标,低阶用扩散模型精确控制动作
- 在杂乱场景中对小物体操作的成功率提升37%,长序列技能组合表现最优
- 适合需要精细操控和复杂推理的机器人应用,如家庭服务或工业装配
尽管端到端的视觉-语言-动作(VLA)模型为机器人操作提供了有前景的范式,但在特定控制数据上微调常会削弱其基础视觉-语言模型(VLM)所具备的深层推理能力。为解决这一根本性权衡,我们提出HiVLA——一种以视觉接地为中心的分层框架,明确将高层语义规划与底层运动控制解耦。高层部分,一个VLM规划器首先进行任务分解与视觉定位,生成包含子任务指令和精确目标边界框的结构化计划。随后,在低层部分引入一种具有新型级联交叉注意力机制的流匹配扩散变换器(DiT)动作专家,以逐步融合全局上下文、高分辨率对象中心裁片及技能语义,使DiT专注于鲁棒执行。该解耦架构在保留VLM零样本推理能力的同时,支持两组件独立优化。大量仿真与真实世界实验表明,HiVLA显著优于现有端到端基线,尤其在长时序技能组合以及杂乱场景中小物体的细粒度操作方面表现突出。
原文摘要 · Abstract (English)
While end-to-end Vision-Language-Action (VLA) models offer a promising paradigm for robotic manipulation, fine-tuning them on narrow control data often compromises the profound reasoning capabilities inherited from their base Vision-Language Models (VLMs). To resolve this fundamental trade-off, we propose HiVLA, a visual-grounded-centric hierarchical framework that explicitly decouples high-level semantic planning from low-level motor control. In high-level part, a VLM planner first performs task decomposition and visual grounding to generate structured plans, comprising a subtask instruction and a precise target bounding box. Then, to translate this plan into physical actions, we introduce a flow-matching Diffusion Transformer (DiT) action expert in low-level part equipped with a novel cascaded cross-attention mechanism. This design sequentially fuses global context, high-resolution object-centric crops and skill semantics, enabling the DiT to focus purely on robust execution. Our decoupled architecture preserves the VLM's zero-shot reasoning while allowing independent improvement of both components. Extensive experiments in simulation and the real world demonstrate that HiVLA significantly outperforms state-of-the-art end-to-end baselines, particularly excelling in long-horizon skill composition and the fine-grained manipulation of small objects in cluttered scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。