arXiv:2608.19490cs.ROcs.CV2026-08

用自生成数据微调视觉语言动作模型,让新机器人同时学会旧技能和新指令。

Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation

论文配图:Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation
图 1 · 摘自论文原文
  • 让零样本模型自动生成交互数据,作为微调的额外训练样本
  • 在真实Aloha机器人上实现多任务泛化,新技能学习效率提升
  • 适合需要快速适配新机器人的具身智能研究者

当前先进的视觉语言动作(VLA)模型如π₀.₅具备强大的语义理解、指令遵循和任务行为能力。然而,在部署到新机器人时,即使硬件配置存在微小差异,也会导致性能显著下降。在新机器人上使用领域内专家数据微调虽能提升特定任务表现,但会损失原有指令遵循能力和行为先验。本文提出一种自监督方法,利用零样本VLA在线生成交互轨迹作为额外训练数据进行微调。实验表明,该方案在目标机器人上可获得强多任务策略:(1)继承零样本模型中提炼的先验任务;(2)保持通用指令遵循能力;(3)通过专家数据学习新技能,且样本效率更高。我们在真实Aloha机器人及RoboTwin新仿真基准上验证了该方法的有效性。视频结果见https://self-supervised-control.pages.dev/

原文摘要 · Abstract (English)

State-of-the-art vision-language-action (VLA) models such as $π_{0.5}$ exhibit strong semantic understanding, instruction following and task behavior. However, when deployed on new robots, even minor mismatches in hardware configuration relative to pretraining can cause severe performance drops. Finetuning the VLA on in-domain expert data from the new embodiment improves performance on the expert task but leads to a loss in its original instruction following and behavioral priors. In this paper, we propose a self-supervised method that generates online interaction rollouts from the zero-shot VLA as additional training data for finetuning. Our experiments show this finetuning scheme yields strong multi-task policies that, on the target robot, (1) inherit prior tasks distilled from the zero-shot model, (2) enable generalist instruction following, while (3) learning new skills from expert data with improved sample efficiency. We demonstrate the success of our approach across test sets probing generalization on a real ALOHA robot and a new simulation benchmark in RoboTwin. Video results are available at https://self-supervised-control.pages.dev/

具身智能多任务自监督机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。