arXiv:2409.17727cs.ROcs.CV2024-09ICRA被引 13

用动作数据微调CLIP,让机器人更好理解语言指令中的动态行为

Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications

  • 在740万帧动作视频上微调CLIP,提升对动态场景的理解能力
  • 在多种语言驱动机器人任务中表现优于现有CLIP模型
  • 适合需要理解动作指令的机器人视觉任务开发者使用

视觉语言模型在各类机器人应用中发挥关键作用。其中,对比图像-文本预训练(CLIP)广泛应用于需同时理解视觉与自然语言的任务。然而,CLIP仅基于静态图像与文本配对训练,尚未充分适配涉及动态动作的机器人任务。本文提出Robotic-CLIP,通过收集并标注大规模动作数据,在309,433个视频(约740万帧)上对CLIP进行对比学习微调。借助动作数据,Robotic-CLIP在继承原版CLIP强大图像表征能力的同时,获得在机器人语境下理解动作的能力。大量实验表明,Robotic-CLIP在多种语言驱动的机器人任务中均优于其他基于CLIP的模型。此外,我们还在真实抓取场景中验证了其实际有效性。

原文摘要 · Abstract (English)

Vision language models have played a key role in extracting meaningful features for various robotic applications. Among these, Contrastive Language-Image Pretraining (CLIP) is widely used in robotic tasks that require both vision and natural language understanding. However, CLIP was trained solely on static images paired with text prompts and has not yet been fully adapted for robotic tasks involving dynamic actions. In this paper, we introduce Robotic-CLIP to enhance robotic perception capabilities. We first gather and label large-scale action data, and then build our Robotic-CLIP by fine-tuning CLIP on 309,433 videos (~7.4 million frames) of action data using contrastive learning. By leveraging action data, Robotic-CLIP inherits CLIP's strong image performance while gaining the ability to understand actions in robotic contexts. Intensive experiments show that our Robotic-CLIP outperforms other CLIP-based models across various language-driven robotic tasks. Additionally, we demonstrate the practical effectiveness of Robotic-CLIP in real-world grasping applications.

机器人视觉CLIP动作理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。