arXiv:2609.05927cs.RO2026-09

用智能体生成可交互的物体组合,提升机器人学习的数据质量

GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning

论文配图:GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning
图 1 · 摘自论文原文
  • 将物体组合生成拆解为解耦重建与相对位姿恢复两步
  • 生成物间碰撞率低于1%,且更匹配结构化指令要求
  • 适合机器人抓取、装配等需要精确物理交互的任务

机器人操作基础模型需要在多样化场景中实现可扩展的评估与数据生成,仿真环境为此提供了可能。自动场景生成具有前景,但以往工作多关注粗粒度布局,忽视细粒度功能型物体组合。针对这一空白,我们提出GIF框架——一种用于生成可交互、功能性物体组合的智能体生成方法。该框架将问题重构为解耦重建与相对位姿恢复。CoGen利用2D与3D生成模型的优势,生成实例解耦的网格并赋予粗略初始位姿;GPRM在几何与物理联合引导下优化相对位姿;视觉语言模型(VLM)验证器则选出最符合结构化规范的候选。我们进一步构建涵盖八类典型接触几何的基准,并与现有先进生成器对比:GIF在资产质量与关系匹配上均有提升,碰撞率降至1%以下。最后,我们用合成数据训练策略,验证了模拟与真实部署中的多样性扩展能力。

原文摘要 · Abstract (English)

Robot manipulation foundation models require scalable evaluation and data generation across diverse scenarios, with simulation providing an environment for both. Automated scene generation offers a promising path, yet prior work has largely emphasized coarse-grained scene layouts rather than fine-grained functional object compositions. Motivated by this gap, we present GIF, an agentic Generation framework for Interactive and Functional object compositions. In this framework, we recast this problem as disentangled reconstruction followed by relative pose recovery. CoGen produces instance-disentangled meshes with coarse initial poses leveraging complementary strengths of 2D and 3D generative models. GPRM refines the relative pose under joint geometric and physical guidance, and a VLM verifier selects the candidate that best matches the structured specification. We further construct a benchmark spanning eight representative contact-geometry classes and compare with state-of-the-art generators; GIF improves both asset quality and relation matching, while reducing collision rate to below 1%. Finally, we synthesize data for policy learning, revealing diversity scaling in both simulation and real-world deployment.

机器人学习物体组合生成模型仿真数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。