arXiv:2511.05680cs.RO2025-11被引 1

用视觉语言模型选技能,让机器人更聪明地组装零件。

VLM-driven Skill Selection for Robotic Assembly Tasks

  • 用视觉语言模型理解任务指令并选择合适操作技能
  • 在多种装配场景中实现高成功率,且技能可解释
  • 适合需要灵活适应新任务的工业自动化场景

本文提出一种结合视觉语言模型(VLMs)与模仿学习的机器人装配框架。系统采用配备夹爪的机器人在三维空间中执行装配操作,融合视觉感知、自然语言理解与已学习的基础技能,实现灵活自适应的机器人操作。实验表明,该方法在多种装配场景中均取得高成功率,同时通过结构化的基础技能分解保持了决策过程的可解释性。

原文摘要 · Abstract (English)

This paper presents a robotic assembly framework that combines Vision-Language Models (VLMs) with imitation learning for assembly manipulation tasks. Our system employs a gripper-equipped robot that moves in 3D space to perform assembly operations. The framework integrates visual perception, natural language understanding, and learned primitive skills to enable flexible and adaptive robotic manipulation. Experimental results demonstrate the effectiveness of our approach in assembly scenarios, achieving high success rates while maintaining interpretability through the structured primitive skill decomposition.

机器人装配视觉语言模型模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。