让机器人理解单个动作,提升视觉语言模型在操作任务中的表现。
RoboAct-CLIP: Video-Driven Pre-training of Atomic Action Understanding for Robotics
- 从机器人视频中提取单一动作,重建纯净训练数据集。
- 通过时间解耦策略,提升模型对原子动作的语义区分能力。
- 适合需要精细动作理解的机器人系统研究者使用。
视觉语言模型(VLMs)已成为机器人系统的关键工具,可通过多模态感知与语义推理实现跨任务泛化、动态环境交互和长时程规划。然而,现有开源VLMs主要针对通用视觉-语言对齐任务训练,难以有效建模机器人操作中关键的时间相关动作语义。当前基于图像的微调方法虽部分适配机器人应用,但忽视视频序列中的时间演化模式,并导致机器人主体、操作物体与环境特征间的视觉混淆,限制了原子动作的语义解耦能力,降低模型泛化性。为此,本文提出RoboAct-CLIP,具双重创新:1)设计语义约束的动作单元分割与重标注框架,重构开源机器人视频数据,生成仅含单一原子动作(如“抓取”)的纯净训练集;2)基于对比语言-图像预训练(CLIP)架构,提出时间解耦微调策略,将视频帧间动作特征与物体中心特征解耦,实现机器人原子动作的层次化表征学习。模拟环境中实验表明,RoboAct-CLIP预训练模型相比基线VLMs成功率提升12%,并在多物体操作任务中展现更优泛化性能。
原文摘要 · Abstract (English)
Visual Language Models (VLMs) have emerged as pivotal tools for robotic systems, enabling cross-task generalization, dynamic environmental interaction, and long-horizon planning through multimodal perception and semantic reasoning. However, existing open-source VLMs predominantly trained for generic vision-language alignment tasks fail to model temporally correlated action semantics that are crucial for robotic manipulation effectively. While current image-based fine-tuning methods partially adapt VLMs to robotic applications, they fundamentally disregard temporal evolution patterns in video sequences and suffer from visual feature entanglement between robotic agents, manipulated objects, and environmental contexts, thereby limiting semantic decoupling capability for atomic actions and compromising model generalizability.To overcome these challenges, this work presents RoboAct-CLIP with dual technical contributions: 1) A dataset reconstruction framework that performs semantic-constrained action unit segmentation and re-annotation on open-source robotic videos, constructing purified training sets containing singular atomic actions (e.g., "grasp"); 2) A temporal-decoupling fine-tuning strategy based on Contrastive Language-Image Pretraining (CLIP) architecture, which disentangles temporal action features across video frames from object-centric characteristics to achieve hierarchical representation learning of robotic atomic actions.Experimental results in simulated environments demonstrate that the RoboAct-CLIP pretrained model achieves a 12% higher success rate than baseline VLMs, along with superior generalization in multi-object manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。