arXiv:2509.18865cs.ROcs.LG2025-09被引 2

用视觉语言融合让机器人一模型搞定多个任务

Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation

  • 融合视觉、语言与机械数据,实现多任务泛化
  • 真实机器人实验成功提升任务成功率
  • 适合需要多场景适应的机器人控制研究者

我们提出基于视觉-语言融合的双边控制模仿学习框架 Bi-VLA,用于动作生成。传统双边控制方法依赖关节角度、速度、力矩和视觉信息进行精确操作,但需为每项任务单独建模,通用性差。Bi-VLA 通过融合领导者-跟随者双边控制中的关节角度、速度、力矩数据,结合 SigLIP 和 FiLM 机制提取的视觉特征与自然语言指令,实现单模型支持多种任务。我们在两类任务上验证:一类需语言补充提示,另一类仅靠视觉区分。真实机器人实验表明,Bi-VLA 能有效理解视觉-语言组合,相比传统方法显著提升任务成功率。结果证实,视觉与语言融合可显著增强系统泛化能力。

原文摘要 · Abstract (English)

We propose Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation (Bi-VLA), a novel framework that extends bilateral control-based imitation learning to handle more than one task within a single model. Conventional bilateral control methods exploit joint angle, velocity, torque, and vision for precise manipulation but require task-specific models, limiting their generality. Bi-VLA overcomes this limitation by utilizing robot joint angle, velocity, and torque data from leader-follower bilateral control with visual features and natural language instructions through SigLIP and FiLM-based fusion. We validated Bi-VLA on two task types: one requiring supplementary language cues and another distinguishable solely by vision. Real-robot experiments showed that Bi-VLA successfully interprets vision-language combinations and improves task success rates compared to conventional bilateral control-based imitation learning. Our Bi-VLA addresses the single-task limitation of prior bilateral approaches and provides empirical evidence that combining vision and language significantly enhances versatility. Experimental results validate the effectiveness of Bi-VLA in real-world tasks. For additional material, please visit the website: https://mertcookimg.github.io/bi-vla/

机器人控制多任务学习视觉语言融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。