让机器人从说明书自动理解零件连接关系,提升装配成功率。
Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals via Vision-Language Models
- 将连接关系作为核心实体建模,构建分层图结构表示装配过程。
- 通过视觉语言模型解析说明书图文,提取超过20种连接类型信息。
- 在家具、玩具等4类真实场景中验证,显著提升复杂装配的执行准确率。
装配的关键在于可靠地形成部件间的连接;然而现有机器人方法在规划装配序列和部件位姿时,常将连接视为次要问题。连接是装配执行的基础物理约束,任务规划虽决定操作顺序,但连接关系的精确建立最终决定装配成败。本文将连接作为装配表示中的显式主实体,为每个装配步骤明确编码连接类型、规格与位置。受人类通过说明书学习装配的启发,我们提出Manual2Skill++,一种基于视觉-语言模型的框架,可自动从装配说明书中提取结构化连接信息。我们将装配任务表示为层次图,节点代表部件与子组件,边显式建模组件间的连接关系。利用大规模视觉-语言模型解析说明书中的符号图示与标注,实例化这些图谱,从而利用人类设计说明中蕴含的丰富连接知识。我们构建了一个包含超过20个装配任务、涵盖多种连接类型的大型数据集,验证了该表示提取方法的有效性,并在模拟环境中评估了从任务理解到执行的完整流程,在家具、玩具及制造组件等四类复杂装配场景中实现与真实世界对应的高精度装配。更多详情请见 https://nus-lins-lab.github.io/Manual2SkillPP/
原文摘要 · Abstract (English)
Assembly hinges on reliably forming connections between parts; yet most robotic approaches plan assembly sequences and part poses while treating connectors as an afterthought. Connections represent the foundational physical constraints of assembly execution; while task planning sequences operations, the precise establishment of these constraints ultimately determines assembly success. In this paper, we treat connections as explicit, primary entities in assembly representation, directly encoding connector types, specifications, and locations for every assembly step. Drawing inspiration from how humans learn assembly tasks through step-by-step instruction manuals, we present Manual2Skill++, a vision-language framework that automatically extracts structured connection information from assembly manuals. We encode assembly tasks as hierarchical graphs where nodes represent parts and sub-assemblies, and edges explicitly model connection relationships between components. A large-scale vision-language model parses symbolic diagrams and annotations in manuals to instantiate these graphs, leveraging the rich connection knowledge embedded in human-designed instructions. We curate a dataset containing over 20 assembly tasks with diverse connector types to validate our representation extraction approach, and evaluate the complete task understanding-to-execution pipeline across four complex assembly scenarios in simulation, spanning furniture, toys, and manufacturing components with real-world correspondence. More detailed information can be found at https://nus-lins-lab.github.io/Manual2SkillPP/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。