arXiv:2602.21157cs.RO2026-02被引 8

HALO让机器人像人一样边想边做,提升复杂任务的推理与泛化能力。

HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning

  • 用分层专家模型实现文本思考、视觉预判和动作预测的统一
  • 在仿真和真实环境中任务成功率比基线高34.1%
  • 适合需要长序列推理与强泛化的机器人任务

视觉-语言-动作(VLA)模型在机器人操作中表现优异,但在长时程或分布外场景下常因缺乏显式多模态推理与世界演化预测机制而失效。现有方法虽引入文本思维链或视觉子目标预测,仍无法提供统一的人类式推理框架。为此,我们提出HALO,一种通过文本任务推理、视觉子目标预测与增强型动作预测的序列化流程,实现具身多模态思维链(EM-CoT)推理的统一VLA模型。采用混合变压器(MoT)架构,将语义推理、视觉预见与动作预测解耦为专用专家,并支持跨专家协作。为实现大规模训练,设计自动化数据合成管道与精细化训练方案。大量实验表明:(1) HALO在仿真与真实环境均表现卓越,在RoboTwin基准上超越基线策略pi_0达34.1%;(2) 训练方案与EM-CoT设计各组件均有效提升任务成功率;(3) 在极端未见环境随机化下展现强泛化能力。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown strong performance in robotic manipulation, but often struggle in long-horizon or out-of-distribution scenarios due to the lack of explicit mechanisms for multimodal reasoning and anticipating how the world will evolve under action. Recent works introduce textual chain-of-thought or visual subgoal prediction within VLA models to reason, but still fail to offer a unified human-like reasoning framework for joint textual reasoning, visual foresight, and action prediction. To this end, we propose HALO, a unified VLA model that enables embodied multimodal chain-of-thought (EM-CoT) reasoning through a sequential process of textual task reasoning, visual subgoal prediction for fine-grained guidance, and EM-CoT-augmented action prediction. We instantiate HALO with a Mixture-of-Transformers (MoT) architecture that decouples semantic reasoning, visual foresight, and action prediction into specialized experts while allowing seamless cross-expert collaboration. To enable HALO learning at scale, we introduce an automated pipeline to synthesize EM-CoT training data along with a carefully crafted training recipe. Extensive experiments demonstrate that: (1) HALO achieves superior performance in both simulated and real-world environments, surpassing baseline policy pi_0 by 34.1% on RoboTwin benchmark; (2) all proposed components of the training recipe and EM-CoT design help improve task success rate; and (3) HALO exhibits strong generalization capabilities under aggressive unseen environmental randomization with our proposed EM-CoT reasoning.

具身智能多模态推理机器人控制思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。