让机器人像人一样持续学新技能,同时不忘记旧技能。
Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation

- 用双时间尺度适配机制拆分短期学习与长期记忆路径。
- 在xArm机器人上实现技能持续扩展,保留率超90%且无需重训。
- 适合需要长期部署的工业机器人场景,兼顾效率与稳定性。
与人类能顺序学习新任务类似,部署在开放环境中的机器人若具备视觉-语言-动作(VLA)模型,也应拥有持续学习能力。然而,现有终身学习模型多聚焦于当前任务表现(可塑性)或旧任务精度保持(稳定性),难以解决二者之间的权衡问题。为此,本文提出一种面向机器人操作的高效缓存终身视觉-语言-动作学习框架(LifelongVLA),通过双时间尺度自适应机制缓解该矛盾,并采用缓存高效的回放策略实现低成本部署。具体而言,设计了一种双时间尺度的LoRA门控模块,将VLA适应分解为短期适配器(增强可塑性)与长期适配器(稳定固化),并通过任务感知门控实现显式调控。在技能回放阶段,提出一种基于随机采样的缓存高效回放策略,避免完整轨迹存储的同时保留更均衡的记忆信号。实验表明,LifelongVLA在xArm机器人上显著优于现有基线,展现出高效的技能拓展能力、超过90%的已学行为保留率,以及对再训练依赖的大幅降低。
原文摘要 · Abstract (English)
Similar to the natural capabilities of humans to sequentially learn new tasks, robots with Vision-Language-Action (VLA) models should possess lifelong learning ability to learn a new task when deployed in open-world environments. However, most recently proposed lifelong learning models aim to effectively learn the current task (plasticity) or maintain high accuracy on previous tasks (stability), while the plasticity-stability trade-off remains largely unsolved in robotic manipulation models. To address this fundamental challenge, we propose a cache-efficient lifelong Vision-Language-Action learning framework for robotic manipulation (i.e., LifelongVLA), which alleviates the plasticity-stability trade-off with a dual-timescale adaptation mechanism while achieving low-cost robotic deployment with a cache-efficient replay strategy. More concretely, we propose a dual-timescale LoRA gating module to decompose VLA adaptation into two lightweight pathways: a short-term adapter for plasticity and a long-term adapter for stable consolidation. These pathways are integrated via a task-aware gate, enabling explicit control of the plasticity-stability trade-off. In the skill replay phase, a cache-efficient stochastic replay strategy is proposed to preserve more balanced retention signals without full-trajectory storage. Finally, experiments show that LifelongVLA outperforms existing baselines, demonstrating efficient skill expansion, robust retention of learned manipulation behaviors, and reduced reliance on retraining for real-world deployment on an xArm robot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。