让视觉语言动作模型主动收集纠错数据,提升学习效率。
RECALL: Recovery Experience Collection for Active Lifelong Learning in Vision-Language-Action Models

- 基于不确定性引导,主动收集需要改进的失误数据。
- 相比被动模仿,新方法使模型适应更快,减少重复训练。
- 适合研究机器人持续学习与高效数据采集的学者。
视觉-语言-动作(VLA)模型通常通过被动模仿学习进行微调,即在策略表现不佳时才收集额外示范。这种方式存在诸多缺陷:需等待机器人失败后才触发数据收集、无法明确指示哪些状态需要监督、且浪费示范者精力在已掌握的任务环节。本文提出一种针对VLA的主动持续学习范式。实验表明,基于不确定性的主动数据收集能显著提升微调效率。但仅使用主动收集的恢复数据微调会导致灾难性遗忘。我们评估了多种持续学习技术,包括基于回放的数据混合和弹性权重固化,发现其在适应新不确定性数据与保留旧知识之间存在权衡。本工作为自回归VLA的主动持续学习提供了实证研究,证明不确定性引导的恢复示范可提高适应效率,同时揭示了向大型机器人策略引入目标新数据时面临的开放挑战。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models are commonly fine-tuned through passive imitation learning, where additional demonstrations are collected for tasks where the policy performs poorly. This approach incurs several downsides: it requires the robot to fail before data collection is triggered, provides little guidance about which states require supervision, and wastes demonstrator effort on redundant parts of the task where the policy already performs well. In this paper, we propose an active, continual learning paradigm for VLAs. We demonstrate that active, uncertainty-guided data collection leads to more efficient fine-tuning than when using passively-collected demonstrations. However, we also find that fine-tuning only on actively-collected recovery data leads to catastrophic forgetting. We evaluate techniques for continual learning, including replay-based data mixing and elastic weight consolidation, and identify tradeoffs between plasticity to uncertainty-guided recovery data and retention of previously learned behaviors. Overall, our work contributes an empirical study of active continual learning for autoregressive VLAs, establishing that uncertainty-guided recovery demonstrations can improve adaptation efficiency while also revealing open challenges when targeted new data is incorporated into large robot policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。