arXiv:2512.02013cs.RO2025-12被引 17

让机器人从目标反推操作步骤,实现自动规划与精准执行

ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation

  • 用多模态手册分步规划动作,结合显式指令与隐式引导
  • 在乐高积木组装任务中成功率提升32%超过现有最先进方法
  • 适合需要长时序规划的复杂机械操作场景

视觉-语言-动作(VLA)模型在机器人场景理解与操作中展现出强大泛化能力。然而,面对需要明确目标状态的长周期任务(如乐高积木拼装或物体重排),现有VLA模型在高层规划与精确操作之间的协调上仍存在挑战。为此,本文提出ManualVLA,一种基于混合变压器(MoT)架构的统一VLA框架,可从目标状态逆向推导出可执行的操作流程。不同于直接将感知输入映射为动作,ManualVLA首先通过规划专家生成包含图像、位置提示和文本指令的多模态手册;随后引入手动思维链(ManualCoT)推理机制,将手册步骤输入动作专家,每个步骤提供显式控制条件,其潜在表示则提供精确操作的隐式指导。为减少数据收集负担,我们开发了基于3D高斯泼溅的高保真数字孪生工具包,可自动生成用于训练规划专家的手册数据。在真实世界测试中,ManualVLA在乐高积木拼装与物体重排任务上的平均成功率比先前的分层最先进基线高出32%。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have recently emerged, demonstrating strong generalization in robotic scene understanding and manipulation. However, when confronted with long-horizon tasks that require defined goal states, such as LEGO assembly or object rearrangement, existing VLA models still face challenges in coordinating high-level planning with precise manipulation. Therefore, we aim to endow a VLA model with the capability to infer the "how" process from the "what" outcomes, transforming goal states into executable procedures. In this paper, we introduce ManualVLA, a unified VLA framework built upon a Mixture-of-Transformers (MoT) architecture, enabling coherent collaboration between multimodal manual generation and action execution. Unlike prior VLA models that directly map sensory inputs to actions, we first equip ManualVLA with a planning expert that generates intermediate manuals consisting of images, position prompts, and textual instructions. Building upon these multimodal manuals, we design a Manual Chain-of-Thought (ManualCoT) reasoning process that feeds them into the action expert, where each manual step provides explicit control conditions, while its latent representation offers implicit guidance for accurate manipulation. To alleviate the burden of data collection, we develop a high-fidelity digital-twin toolkit based on 3D Gaussian Splatting, which automatically generates manual data for planning expert training. ManualVLA demonstrates strong real-world performance, achieving an average success rate 32% higher than the previous hierarchical SOTA baseline on LEGO assembly and object rearrangement tasks.

机器人操作多模态规划长程任务视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。