自动生成组装说明书,让AI学会拆解并重建复杂乐高积木。
Learning to Build by Building Your Own Instructions
- AI通过自拆乐高并拍照生成步骤说明书,构建显式记忆。
- 可处理平均31块积木、需百步操作的大型乐高模型。
- 适合研究视觉结构理解与具身智能的科研人员。
复杂视觉对象的结构理解是人工智能中尚未解决的重要问题。为研究此问题,我们提出一种针对LTRON中最新提出的“拆解与重建”任务的新方法,该任务要求智能体在单次交互会话中学习未见过的乐高积木组装结构。我们开发了名为 extbf{ ous}的智能体,能够自主生成视觉说明书:通过拆解未知积木并定期保存图像,构建出重建所需的步骤序列。这些说明形成显式记忆,使模型可逐步推理,无需依赖长期隐式记忆,从而支持训练更大规模的乐高模型。为验证该方法,我们发布了一个新数据集,包含程序生成的乐高车辆,每辆平均含31块积木,拆装需超过100步。模型通过在线模仿学习进行训练,能从自身错误中学习。此外,我们还对LTRON和拆建任务进行了若干改进,简化环境并提升可用性。
原文摘要 · Abstract (English)
Structural understanding of complex visual objects is an important unsolved component of artificial intelligence. To study this, we develop a new technique for the recently proposed Break-and-Make problem in LTRON where an agent must learn to build a previously unseen LEGO assembly using a single interactive session to gather information about its components and their structure. We attack this problem by building an agent that we call \textbf{\ours} that is able to make its own visual instruction book. By disassembling an unseen assembly and periodically saving images of it, the agent is able to create a set of instructions so that it has the information necessary to rebuild it. These instructions form an explicit memory that allows the model to reason about the assembly process one step at a time, avoiding the need for long-term implicit memory. This in turn allows us to train on much larger LEGO assemblies than has been possible in the past. To demonstrate the power of this model, we release a new dataset of procedurally built LEGO vehicles that contain an average of 31 bricks each and require over one hundred steps to disassemble and reassemble. We train these models using online imitation learning which allows the model to learn from its own mistakes. Finally, we also provide some small improvements to LTRON and the Break-and-Make problem that simplify the learning environment and improve usability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。