让机器人模型与专家协作,双向提升操作能力与学习效率。
VLA Model-Expert Collaboration for Bi-directional Manipulation Learning
- 用少量专家动作指导VLA模型,降低人工负担
- 多任务实验显示成功率显著提升,跨任务泛化增强
- 适合人机协同机器人系统研究者与开发者
视觉-语言-动作(VLA)模型的兴起催生了机器人操作领域的基础模型。尽管已有显著进展,其在多任务操作中的泛化能力仍有限。本文提出一种VLA模型与专家协同框架,仅需少量专家动作即可提升模型性能,既减轻人工负担,又增强模型可靠性与泛化能力。协作过程中收集的操作数据可进一步优化VLA模型,同时人类参与者技能亦得到提升,形成双向学习闭环。多组实验验证了该系统在不同VLA模型上的有效性,任务成功率普遍提高。此外,通过脑机接口(BCI)验证,该系统在低速动作场景中显著提升了操作效率。这些成果为机器人领域基础模型时代的人机交互发展提供了新路径。
原文摘要 · Abstract (English)
The emergence of vision-language-action (VLA) models has given rise to foundation models for robot manipulation. Although these models have achieved significant improvements, their generalization in multi-task manipulation remains limited. This study proposes a VLA model-expert collaboration framework that leverages a limited number of expert actions to enhance VLA model performance. This approach reduces expert workload relative to manual operation while simultaneously improving the reliability and generalization of VLA models. Furthermore, manipulation data collected during collaboration can further refine the VLA model, while human participants concurrently enhance their skills. This bi-directional learning loop boosts the overall performance of the collaboration system. Experimental results across various VLA models demonstrate the effectiveness of the proposed system in collaborative manipulation and learning, as evidenced by improved success rates across tasks. Additionally, validation using a brain-computer interface (BCI) indicates that the collaboration system enhances the efficiency of low-speed action systems by involving VLA model during manipulation. These promising results pave the way for advancing human-robot interaction in the era of foundation models for robotics. (Project website: https://aoqunjin.github.io/Expert-VLA/)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。