arXiv:2603.08519cs.RO2026-03被引 12

通过分步建模提升机器人长程操作的鲁棒性,成功率高达97%

AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models

  • 用大模型拆解任务为原子子步骤,实现细粒度控制
  • 在潜空间中用世界模型评分动作块,减少错误累积
  • 无需真实机器人试错,适合长程复杂任务研究

视觉-语言-动作(VLA)模型在通用机器人操作中展现出巨大潜力。现有方法多依赖高层指令进行监督微调,缺乏中间阶段引导,导致长程任务中错误不断累积。为此,我们提出AtomVLA,首个结合可扩展离线后训练流程的子任务感知VLA框架。该框架利用大语言模型将高阶演示分解为细粒度原子子任务,并借助预训练的预测世界模型,在潜空间中对候选动作片段与子任务目标进行评分,有效缓解误差传播,显著提升长程任务鲁棒性。此外,该方法支持高效的组相对策略优化,避免了物理机器人在线采样的高昂成本。大量仿真验证表明,AtomVLA在扰动下仍保持强鲁棒性;在LIBERO基准上平均成功率达97.0%,在LIBERO-PRO上达48.0%。真实世界实验使用Galaxea R1 Lite平台,验证了其在多样化任务尤其是长程任务中的广泛适用性。所有数据集、模型权重和代码将在论文接收后公开。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The execution of complex multi-step behaviors in VLA models can be improved by robust instruction grounding, a critical component for effective control. However, current paradigms predominantly rely on coarse, high-level task instructions during supervised fine-tuning. This instruction grounding gap leaves models without explicit intermediate guidance, leading to severe compounding errors in long-horizon tasks. Therefore, bridging this instruction gap and providing scalable post-training for VLA models is urgent. To tackle this problem, we propose \method, the first subtask-aware VLA framework integrated with a scalable offline post-training pipeline. Our framework leverages a large language model to decompose high-level demonstrations into fine-grained atomic subtasks. This approach utilizes a pretrained predictive world model to score candidate action chunks against subtask goals in the latent space, mitigating error accumulation while significantly improving long-horizon robustness. Furthermore, this approach enables highly efficient Group Relative Policy Optimization without the prohibitive expenses associated with online rollouts on physical robots. Extensive simulations validate that our AtomVLA maintains strong robustness under perturbations. When evaluated against fundamental baseline models, it achieves an average success rate of 97.0\% on the LIBERO benchmark and 48.0\% on the LIBERO-PRO benchmark. Finally, experiments conducted in the real world using the Galaxea R1 Lite platform confirm its broad applicability across diverse tasks, especially long-horizon tasks. All datasets, checkpoints, and code will be released to the public domain following the acceptance of this work for future research.

机器人操作长程任务世界模型后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。