arXiv:2609.06279cs.ROcs.CV2026-09

用图像编辑生成可执行的机器人操作数据,兼顾语义与物理可行性。

IM-ENGINE: Image Editing for Embodied Data Generation

论文配图:IM-ENGINE: Image Editing for Embodied Data Generation
图 1 · 摘自论文原文
  • 以图像编辑为中间表示,结合模拟器先验重建3D状态。
  • 生成的抓取和目标状态经物理验证,可直接用于机器人学习。
  • 适合需要高质量标注数据的机器人操控任务研究者。

基于学习的操纵需要兼具语义意义和物理可执行性的监督,但现有数据管道通常只提供其中一种。人类示范虽能体现意图,但采集成本高且受限于人机体态差异;仿真可大规模生成数据,但常缺乏功能性描述。本文提出IM-ENGINE,一种基于模拟器的实体数据生成流水线,利用图像编辑作为中间表示。给定包含已知几何、深度、分割和相机参数的渲染场景,IM-ENGINE在图像空间注入任务相关语义,借助模拟器先验和不变锚物体恢复显式3D状态,进行物理优化,并转化为机器人可执行的监督信号。该流程应用于灵巧抓取合成与目标状态生成:抓取任务中,生成人手抓取图像,恢复手物交互,重定向至机器人手并优化为物理有效的抓取;目标生成中,将渲染场景编辑为期望结果,恢复目标物体姿态,生成物理合理且语义明确的目标与轨迹。该方法结合生成语义先验与模拟器约束,实现了可扩展的任务相关监督。

原文摘要 · Abstract (English)

Learning-based manipulation requires supervision that is both semantically meaningful and physically executable, but current data pipelines often provide only one of these properties. Human demonstrations capture intent but are costly to collect and constrained by the human-robot embodiment gap, while simulation can scale data generation but often under-specifies functional behavior. We present IM-ENGINE, a simulator-grounded pipeline that uses image editing as an intermediate representation for embodied data generation. Given a rendered scene with known geometry, depth, segmentation, and camera parameters, IM-ENGINE edits the image to inject task-relevant semantics, recovers explicit 3D state using simulator priors and an unchanged anchor object, refines the state in physics, and converts it into robot-executable supervision. We instantiate the pipeline for dexterous grasp synthesis and goal-state generation. For grasping, IM-ENGINE generates a human grasp in image space, recovers the hand-object interaction, retargets it to a robot hand, and refines it into physically validated robot grasps. For goal generation, it edits a rendered scene into a desired outcome, recovers the target-object pose, and refines it into physically valid, semantically meaningful goals and trajectories. This combination of generative semantic priors and simulator grounding enables scalable task-relevant supervision for robot learning.

机器人学习图像编辑模拟器抓取生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。