arXiv:2509.05314cs.ROcs.AI2025-09被引 18

用3D轨迹生成更真实的机器人操作视频,减少人工干预。

ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory

  • 结合3D占用图与优化轨迹规划,生成真实可行的机械臂路径。
  • 相比2D方法,视频视觉质量显著提升,轨迹更符合物理约束。
  • 适合机器人仿真、具身智能研究者快速生成高质量操作视频。

机器人操作数据稀缺仍是重大挑战。尽管扩散模型为生成操作视频提供了新思路,但现有方法多依赖2D轨迹,存在3D空间模糊问题。本文提出ManipDreamer3D框架,从输入图像和文本指令生成具备3D感知的机器人操作视频。首先,从第三人称视角重建3D占用表示,并规划最小路径长度且避障的3D末端执行器轨迹;随后,利用潜在空间编辑技术,结合初始图像隐向量与优化后的3D轨迹,驱动专门训练的轨迹到视频扩散模型生成抓取-放置类视频。该方法实现了自主规划的合理3D轨迹,大幅降低人工干预需求。实验表明,生成视频在视觉质量上优于现有方法。

原文摘要 · Abstract (English)

Data scarcity continues to be a major challenge in the field of robotic manipulation. Although diffusion models provide a promising solution for generating robotic manipulation videos, existing methods largely depend on 2D trajectories, which inherently face issues with 3D spatial ambiguity. In this work, we present a novel framework named ManipDreamer3D for generating plausible 3D-aware robotic manipulation videos from the input image and the text instruction. Our method combines 3D trajectory planning with a reconstructed 3D occupancy map created from a third-person perspective, along with a novel trajectory-to-video diffusion model. Specifically, ManipDreamer3D first reconstructs the 3D occupancy representation from the input image and then computes an optimized 3D end-effector trajectory, minimizing path length while avoiding collisions. Next, we employ a latent editing technique to create video sequences from the initial image latent and the optimized 3D trajectory. This process conditions our specially trained trajectory-to-video diffusion model to produce robotic pick-and-place videos. Our method generates robotic videos with autonomously planned plausible 3D trajectories, significantly reducing human intervention requirements. Experimental results demonstrate superior visual quality compared to existing methods.

机器人操作3D生成扩散模型轨迹规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。