arXiv:2410.15536cs.ROcs.AI2024-10CVPR被引 18

用真实图像生成可训练机器人的仿真任务,让虚拟机器人更快上手真实场景。

GRS: Generating Robotic Simulation Tasks from Real-World Images

  • 通过视觉语言模型理解真实图像中的物体与场景
  • 自动匹配仿真资源并生成可执行的任务环境
  • 用迭代优化的路由机制确保仿真与任务精准对齐

我们提出GRS(Generating Robotic Simulation tasks),一种解决真实世界到仿真转换的系统。GRS基于单张RGB-D图像生成可执行任务的数字孪生仿真环境,用于虚拟代理训练。该系统采用视觉语言模型(VLMs)分三阶段运行:1)使用SAM2进行分割与物体描述,实现场景理解;2)将真实物体匹配至仿真可用资产;3)生成合适任务。通过生成的测试套件确保仿真与任务的一致性,并引入路由器机制,迭代优化仿真与测试代码。实验表明,该系统在物体对应与任务环境生成方面表现有效,其新型路由器机制显著提升对齐精度。

原文摘要 · Abstract (English)

We introduce GRS (Generating Robotic Simulation tasks), a system addressing real-to-sim for robotic simulations. GRS creates digital twin simulations from single RGB-D observations with solvable tasks for virtual agent training. Using vision-language models (VLMs), our pipeline operates in three stages: 1) scene comprehension with SAM2 for segmentation and object description, 2) matching objects with simulation-ready assets, and 3) generating appropriate tasks. We ensure simulation-task alignment through generated test suites and introduce a router that iteratively refines both simulation and test code. Experiments demonstrate our system's effectiveness in object correspondence and task environment generation through our novel router mechanism.

机器人仿真视觉语言模型数字孪生任务生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。