arXiv:2603.07648cs.ROcs.AI2026-03中稿 · CVPR被引 12

提出原子化技能学习框架,让机器人能持续学习复杂任务。

AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots

  • 用专家混合模型构建可扩展的原子技能库,支持精细动作生成。
  • 在仿真中比基线提升10%,真实世界任务成功率高出21%。
  • 适合需要长期学习和多步规划的机器人场景。

视觉-语言-动作(VLA)模型在机器人操作中展现出潜力,但现实任务常需长时序、多步骤求解与持续技能积累,现有模型因使用整体动作解码器且训练数据聚合,导致可扩展性差。为此,我们提出AtomicVLA,一个统一的规划-执行框架,联合生成任务级计划、原子技能抽象与细粒度动作。通过技能引导的专家混合(SG-MoE),每个专家专注掌握通用而精确的原子技能,构建可扩展的技能库。引入灵活路由编码器,自动为新技能分配专属专家,实现持续学习。实验表明,在仿真中,AtomicVLA在LIBERO上优于π₀ 2.4%,在LIBERO-LONG上提升10%,在CALVIN上平均任务长度分别优于π₀和π₀.₅ 0.22和0.25。真实世界长时序任务中,性能超越基线18.3%,持续学习任务提升21%。结果验证了原子技能抽象与动态专家组合在长时序与终身学习任务中的有效性。

原文摘要 · Abstract (English)

Recent advances in Visual-Language-Action (VLA) models have shown promising potential for robotic manipulation tasks. However, real-world robotic tasks often involve long-horizon, multi-step problem-solving and require generalization for continual skill acquisition, extending beyond single actions or skills. These challenges present significant barriers for existing VLA models, which use monolithic action decoders trained on aggregated data, resulting in poor scalability. To address these challenges, we propose AtomicVLA, a unified planning-and-execution framework that jointly generates task-level plans, atomic skill abstractions, and fine-grained actions. AtomicVLA constructs a scalable atomic skill library through a Skill-Guided Mixture-of-Experts (SG-MoE), where each expert specializes in mastering generic yet precise atomic skills. Furthermore, we introduce a flexible routing encoder that automatically assigns dedicated atomic experts to new skills, enabling continual learning. We validate our approach through extensive experiments. In simulation, AtomicVLA outperforms $π_{0}$ by 2.4\% on LIBERO, 10\% on LIBERO-LONG, and outperforms $π_{0}$ and $π_{0.5}$ by 0.22 and 0.25 in average task length on CALVIN. Additionally, our AtomicVLA consistently surpasses baselines by 18.3\% and 21\% in real-world long-horizon tasks and continual learning. These results highlight the effectiveness of atomic skill abstraction and dynamic expert composition for long-horizon and lifelong robotic tasks. The project page is \href{https://zhanglk9.github.io/atomicvla-web/}{here}.

机器人技能学习持续学习多步规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。