用说明书图解指导3D家具零件组装,提升精度与泛化能力
Manual-PA: Learning 3D Part Assembly from Instruction Diagrams
- 基于图解分阶段求解:先语义对齐零件与图示,再推断6D姿态
- 在PartNet上装配成功率超现有方法,真实IKEA家具也表现良好
- 适合需要从图文引导中学习复杂组装的机器人与AI系统
家具组装本质上是离散-连续优化问题,需选择零件并估计其物理合理的连接位姿。由于解空间组合爆炸且稀疏,当前机器学习模型难以有效学习。本文提出Manual-PA,一种基于Transformer的指令图引导3D零件组装框架。利用随附的图解说明书,将任务分解为离散与连续两阶段:首先通过对比学习使3D零件与说明书图示语义对齐,预测装配顺序;再根据最终成品图推断每个零件的6D位姿。在PartNet基准数据集上验证,结合图解与零件顺序显著提升装配性能。此外,Manual-PA在真实IKEA家具组装数据集(IKEA-Manual)上展现出强泛化能力。
原文摘要 · Abstract (English)
Assembling furniture amounts to solving the discrete-continuous optimization task of selecting the furniture parts to assemble and estimating their connecting poses in a physically realistic manner. The problem is hampered by its combinatorially large yet sparse solution space thus making learning to assemble a challenging task for current machine learning models. In this paper, we attempt to solve this task by leveraging the assembly instructions provided in diagrammatic manuals that typically accompany the furniture parts. Our key insight is to use the cues in these diagrams to split the problem into discrete and continuous phases. Specifically, we present Manual-PA, a transformer-based instruction Manual-guided 3D Part Assembly framework that learns to semantically align 3D parts with their illustrations in the manuals using a contrastive learning backbone towards predicting the assembly order and infers the 6D pose of each part via relating it to the final furniture depicted in the manual. To validate the efficacy of our method, we conduct experiments on the benchmark PartNet dataset. Our results show that using the diagrams and the order of the parts lead to significant improvements in assembly performance against the state of the art. Further, Manual-PA demonstrates strong generalization to real-world IKEA furniture assembly on the IKEA-Manual dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。