让机器人看懂家具说明书,自动完成组装。
Manual2Skill: Learning to Read Manuals and Acquire Robotic Skills for Furniture Assembly Using Vision-Language Models
- 用视觉语言模型解析说明书图文,构建分层装配图谱
- 通过位姿估计和路径规划实现真实机械臂的精准操作
- 可处理长时序复杂任务,适合家庭服务机器人
人类能通过理解抽象说明书完成复杂装配任务,而机器人仍面临巨大挑战。本文提出Manual2Skill框架,利用视觉语言模型(VLM)从说明书图像中提取结构化信息,构建包含零件、子部件及其关系的层次化装配图。通过位姿估计模型预测各步骤组件间的6D相对位姿,结合运动规划模块生成可执行的机器人动作序列。在多个真实IKEA家具组装任务中验证了该方法的有效性,展现了对长时序操纵任务的高效与精确处理能力,显著提升了机器人从说明书学习并执行复杂操作的实用性。该研究推动了机器人向类人理解与执行能力迈进。
原文摘要 · Abstract (English)
Humans possess an extraordinary ability to understand and execute complex manipulation tasks by interpreting abstract instruction manuals. For robots, however, this capability remains a substantial challenge, as they cannot interpret abstract instructions and translate them into executable actions. In this paper, we present Manual2Skill, a novel framework that enables robots to perform complex assembly tasks guided by high-level manual instructions. Our approach leverages a Vision-Language Model (VLM) to extract structured information from instructional images and then uses this information to construct hierarchical assembly graphs. These graphs represent parts, subassemblies, and the relationships between them. To facilitate task execution, a pose estimation model predicts the relative 6D poses of components at each assembly step. At the same time, a motion planning module generates actionable sequences for real-world robotic implementation. We demonstrate the effectiveness of Manual2Skill by successfully assembling several real-world IKEA furniture items. This application highlights its ability to manage long-horizon manipulation tasks with both efficiency and precision, significantly enhancing the practicality of robot learning from instruction manuals. This work marks a step forward in advancing robotic systems capable of understanding and executing complex manipulation tasks in a manner akin to human capabilities.Project Page: https://owensun2004.github.io/Furniture-Assembly-Web/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。