实测发现视觉语言动作模型能持续学新技能不丢旧本领
Can Vision-Language-Action Models Learn from Real-World Data Continually without Forgetting?

- 用真实机器人任务构建持续学习基准,测试10个不同操作技能
- 经验回放方法显著减少遗忘,比直接微调效果好得多
- 适合长期运行的机器人系统开发者参考
视觉-语言-动作(VLA)模型为通用机器人提供了前景,但其真实部署需具备在不遗忘旧技能的前提下持续学习新技能的能力。尽管已有研究在模拟环境中探索了VLA模型的持续学习,但在真实物理条件下的挑战仍鲜有研究。为此,我们构建了一个包含十种不同序列操作任务的真实世界持续学习基准,涵盖单臂和双臂配置。通过大量实验发现,简单的顺序微调会导致严重灾难性遗忘;而经过良好配置的经验回放(ER)方法可有效缓解遗忘,且在同等计算预算下优于联合多任务训练。值得注意的是,结合实证结果,我们成功实现了在完整10任务异构流上的稳定持续学习,既保留已有能力,又能适应多样新技能。本工作基于真实世界实验,为部署鲁棒、长寿命机器人策略提供了可操作洞见。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models provide a promising foundation for general-purpose robotics, yet their real-world deployment demands the ability to continually acquire new skills without forgetting prior ones. While recent studies have explored continual learning for VLA models in simulated settings, the challenge remains largely unexamined under realistic physical conditions. To bridge this gap, we construct a real-world continual learning benchmark comprising ten diverse sequential manipulation tasks across both single-arm and bimanual configurations. Through extensive experiments on this benchmark, we find that naive sequential fine-tuning leads to severe catastrophic forgetting, whereas a well-configured experience replay (ER) approach can effectively mitigate forgetting and outperform joint multi-task training under equivalent computational budgets. Notably, by synthesizing our empirical findings, we successfully achieve stable continual learning across the full 10-task heterogeneous stream, retaining previously acquired capabilities while adapting to diverse new skills in real-world deployment. This work presents an empirical study grounded in real-world continual VLA learning and offers actionable insights for deploying robust, long-lived robotic policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。