arXiv:2512.14217cs.CVcs.RO2025-12被引 3

用深度轨迹生成可控机器人演示视频,提升真实场景下的操作成功率。

DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos

  • 从轨迹中提取深度、语义、形状和运动等多维特征注入扩散模型。
  • 联合生成对齐的彩色与深度视频,提升时空一致性,成功率达78.3%。
  • 适合需要高保真机器人演示生成的研究者或工业应用开发者。

视频扩散模型为具身智能提供了强大的现实世界模拟器,但在机器人操作中的可控性仍受限。现有基于轨迹的视频生成方法多依赖二维轨迹或单一模态输入,难以生成可控且一致的机器人演示。我们提出DRAW2ACT,一种深度感知的轨迹条件视频生成框架,从输入轨迹中提取深度、语义、形状和运动等多正交表征,并将其注入扩散模型。此外,我们提出联合生成空间对齐的RGB与深度视频,利用跨模态注意力机制和深度监督增强时空一致性。最后,设计一个多模态策略模型,基于生成的RGB与深度序列回归机器人关节角。在Bridge V2、Berkeley Autolab及仿真基准上的实验表明,DRAW2ACT在视觉保真度与一致性上均优于现有基线,操作成功率提升至78.3%。

原文摘要 · Abstract (English)

Video diffusion models provide powerful real-world simulators for embodied AI but remain limited in controllability for robotic manipulation. Recent works on trajectory-conditioned video generation address this gap but often rely on 2D trajectories or single modality conditioning, which restricts their ability to produce controllable and consistent robotic demonstrations. We present DRAW2ACT, a depth-aware trajectory-conditioned video generation framework that extracts multiple orthogonal representations from the input trajectory, capturing depth, semantics, shape and motion, and injects them into the diffusion model. Moreover, we propose to jointly generate spatially aligned RGB and depth videos, leveraging cross-modality attention mechanisms and depth supervision to enhance the spatio-temporal consistency. Finally, we introduce a multimodal policy model conditioned on the generated RGB and depth sequences to regress the robot's joint angles. Experiments on Bridge V2, Berkeley Autolab, and simulation benchmarks show that DRAW2ACT achieves superior visual fidelity and consistency while yielding higher manipulation success rates compared to existing baselines.

视频生成扩散模型机器人控制多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。