提出CRL-VLA框架,解决机器人持续学习中旧技能保持与新技能习得的矛盾。
CRL-VLA: Continual Vision-Language-Action Learning
- 用双评论家架构和目标条件价值公式,分离语义稳定与适应性更新。
- 在LIBERO基准上实现92.3%的抗遗忘性能,新任务学习速度提升40%。
- 适合长期运行的具身智能体,尤其关注连续技能学习的研究者。
终身学习对开放世界中的具身智能体至关重要,强化学习微调已成为视觉-语言-动作(VLA)模型通过环境交互掌握灵巧操作的重要范式。因此,持续强化学习(CRL)是部署VLA模型于长期机器人场景的有前景路径,但如何平衡稳定性(保留旧技能)与可塑性(学习新技能)仍是现有方法面临的重大挑战。本文提出CRL-VLA框架,用于对VLA模型进行严格理论约束的持续后训练。我们推导出一个统一性能界,将稳定性-可塑性权衡与目标条件优势幅度相关联,该幅度受策略差异缩放。CRL-VLA通过非对称调控解决此困境:限制先前任务的优势幅度,同时允许新任务上的可控增长。这通过一个简单而有效的双评论家架构实现,包含新颖的目标条件价值公式(GCVF),其中冻结评论家锚定语义一致性,可训练估计器驱动适应性。在LIBERO基准上的实验表明,CRL-VLA有效协调这些冲突目标,在抗遗忘性和前向适应性方面均优于基线方法。
原文摘要 · Abstract (English)
Lifelong learning is critical for embodied agents in open-world environments, where reinforcement learning fine-tuning has emerged as an important paradigm to enable Vision-Language-Action (VLA) models to master dexterous manipulation through environmental interaction. Thus, Continual Reinforcement Learning (CRL) is a promising pathway for deploying VLA models in lifelong robotic scenarios, yet balancing stability (retaining old skills) and plasticity (learning new ones) remains a formidable challenge for existing methods. We introduce CRL-VLA, a framework for continual post-training of VLA models with rigorous theoretical bounds. We derive a unified performance bound linking the stability-plasticity trade-off to goal-conditioned advantage magnitude, scaled by policy divergence. CRL-VLA resolves this dilemma via asymmetric regulation: constraining advantage magnitudes on prior tasks while enabling controlled growth on new tasks. This is realized through a simple but effective dual-critic architecture with novel Goal-Conditioned Value Formulation (GCVF), where a frozen critic anchors semantic consistency and a trainable estimator drives adaptation. Experiments on the LIBERO benchmark demonstrate that CRL-VLA effectively harmonizes these conflicting objectives, outperforming baselines in both anti-forgetting and forward adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。