无需增加参数,让视觉语言动作模型持续学习新技能。
Continually Evolving Skill Knowledge in Vision Language Action Model
- 用知识驱动框架实现无新增参数的持续模仿学习。
- 仅用1%数据重放,在LIBERO上超越多数基线模型。
- 适合需要长期积累技能的机器人系统开发者。
视觉-语言-动作(VLA)模型在预训练中展现良好知识累积能力,但持续学习仍具挑战,尤其在高效适应方面。现有持续模仿学习(CIL)方法常依赖额外参数或外部模块,限制了大VLA模型的可扩展性。我们提出Stellar VLA,一种不增加网络参数的知识驱动型CIL框架。设计了两种渐进式扩展变体:T-Stellar用于任务中心建模,TS-Stellar用于分层任务-技能结构。Stellar VLA通过联合优化任务表征与学习到的知识空间,实现自演化知识学习。提出基于知识关联与Top-K语义嵌入的知识引导专家路由机制,实现任务专精而不增加模型规模。在LIBERO基准测试中,Stellar VLA在仅使用1%数据重放的情况下,性能优于众多VLA与CIL基线。在双臂平台的真实世界评估中,验证了跨不同实体和场景配置的有效知识迁移。TS-Stellar在分层操作任务中表现优异,可视化结果揭示出稳健的知识保留与任务发现能力。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models show promising knowledge accumulation ability from pretraining, yet continual learning in VLA remains challenging, especially for efficient adaptation. Existing continual imitation learning (CIL) methods often rely on additional parameters or external modules, limiting scalability for large VLA models. We propose Stellar VLA, a knowledge-driven CIL framework without increasing network parameters. Two progressively extended variants are designed: T-Stellar for flat task-centric modeling and TS-Stellar for hierarchical task-skill structure. Stellar VLA enables self-evolving knowledge learning by jointly optimizing task representations and a learned knowledge space. We propose a knowledge-guided expert routing mechanism conditioned on knowledge relation and Top-K semantic embeddings, enabling task specialization without increasing model size. Experiments on the LIBERO benchmark show that Stellar VLAs achieve strong performance among both VLA and CIL baselines, using only 1 % data replay. Real-world evaluation on a dual-arm platform with distinct embodiment and scene configurations validates effective knowledge transfer. TS-Stellar excels in hierarchical manipulation, and visualizations reveal robust knowledge retention and task discovery. Project Website: https://stellarvla.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。