用AI自动生成可被机器人组装的积木作品,全程无需人工干预。
Blox-Net: Generative Design-for-Robot-Assembly Using VLM Supervision, Physics Simulation, and a Robot with Reset
- 结合视觉语言模型与物理仿真,从文字指令生成可组装结构。
- 设计识别准确率达63.5%,经重设计后10次连续组装成功率接近100%。
- 适合对自动化装配、生成式AI落地感兴趣的开发者与研究者。
生成式AI在文本、代码和图像生成方面表现出色。受工业领域“面向装配的设计”研究启发,我们提出新问题:生成式面向机器人装配(GDfRA)。任务是根据自然语言提示(如“长颈鹿”)和可用物理组件(如3D打印积木)的图像,生成一个装配体——即这些组件的空间布局及机器人构建指令。输出需满足两点:1)与请求对象相似;2)能由六自由度机械臂配合吸盘可靠完成装配。为此,我们提出Blox-Net系统,融合生成式视觉语言模型、计算机视觉、物理仿真、扰动分析、运动规划与真实机器人实验,实现一类GDfRA问题的最小人工监督求解。Blox-Net在“可识别性”上达到63.5%的Top-1准确率(由视觉语言模型评估,如长颈鹿识别)。经自动扰动重设计后,机器人在10次连续装配中几乎全成功,仅需在每次装配前人工重启。令人惊讶的是,从文字输入(如“长颈鹿”)到可靠物理装配,整个过程实现了零人工干预。
原文摘要 · Abstract (English)
Generative AI systems have shown impressive capabilities in creating text, code, and images. Inspired by the rich history of research in industrial ''Design for Assembly'', we introduce a novel problem: Generative Design-for-Robot-Assembly (GDfRA). The task is to generate an assembly based on a natural language prompt (e.g., ''giraffe'') and an image of available physical components, such as 3D-printed blocks. The output is an assembly, a spatial arrangement of these components, and instructions for a robot to build this assembly. The output must 1) resemble the requested object and 2) be reliably assembled by a 6 DoF robot arm with a suction gripper. We then present Blox-Net, a GDfRA system that combines generative vision language models with well-established methods in computer vision, simulation, perturbation analysis, motion planning, and physical robot experimentation to solve a class of GDfRA problems with minimal human supervision. Blox-Net achieved a Top-1 accuracy of 63.5% in the ''recognizability'' of its designed assemblies (eg, resembling giraffe as judged by a VLM). These designs, after automated perturbation redesign, were reliably assembled by a robot, achieving near-perfect success across 10 consecutive assembly iterations with human intervention only during reset prior to assembly. Surprisingly, this entire design process from textual word (''giraffe'') to reliable physical assembly is performed with zero human intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。