轻量级视觉语言模型让机器人快速决策且记性好。
Towards Fast, Memory-based and Data-Efficient Vision-Language Policy
- 用10亿参数大模型微调小数据,构建轻量记忆型策略
- 推理速度更快,长任务表现比基线高18.8%
- 适合资源有限但需高效学习的机器人应用
视觉语言模型(VLM)在互联网规模数据上预训练后展现出向机器人学习迁移的潜力。然而,现有方法面临三大挑战:(1)大规模参数导致昂贵的推理成本;(2)因数据模态不匹配引发频繁领域偏移;(3)难以处理过往或未来经验。本文提出LiteVLP,一种轻量、基于记忆、通用的视觉语言策略生成模型。LiteVLP基于一个10亿参数的预训练VLM,仅在小规模对话式机器人数据集上微调。大量实验表明,LiteVLP在VIMA-Bench上优于当前最优视觉语言策略,且训练时间极短。此外,其推理速度显著提升,同时保持高精度。在长时序操作任务中,其记忆能力远超最佳基线模型,性能提升18.8%。这些结果表明,LiteVLP是将VLM智能融入机器人学习的有力候选。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) pretrained on Internet-scale vision-language data have demonstrated the potential to transfer their knowledge to robotic learning. However, the existing paradigm encounters three critical challenges: (1) expensive inference cost resulting from large-scale model parameters, (2) frequent domain shifts caused by mismatched data modalities, and (3) limited capacity to handle past or future experiences. In this work, we propose LiteVLP, a lightweight, memory-based, and general-purpose vision-language policy generation model. LiteVLP is built upon a pre-trained 1B-parameter VLM and fine-tuned on a tiny-scale and conversation-style robotic dataset. Through extensive experiments, we demonstrate that LiteVLP outperforms state-of-the-art vision-language policy on VIMA-Bench, with minimal training time. Furthermore, LiteVLP exhibits superior inference speed while maintaining exceptional high accuracy. In long-horizon manipulation tasks, LiteVLP also shows remarkable memory ability, outperforming the best-performing baseline model by 18.8%. These results highlight LiteVLP as a promising model to integrating the intelligence of VLMs into robotic learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。