arXiv:2507.06710cs.RO2025-07ICCV被引 9

用时空感知扩散模型提升机器人视觉模仿学习的3D与4D理解能力。

Spatial-Temporal Aware Visuomotor Diffusion Policy Learning

  • 基于动态高斯世界模型,从单视角RGB-D图构建3D场景并预测未来场景。
  • 在17个仿真任务和3个真实机器人任务中,平均成功率提升6.45%~16.4%。
  • 适合需要精准空间结构与时间关系建模的复杂机器人操作任务。

视觉模仿学习能有效让机器人掌握多样化任务,但现有方法多依赖监督轨迹克隆,缺乏对3D空间与4D时空关系的充分感知,难以满足真实部署需求。本文提出4D扩散策略(DP4),将时空感知融入基于扩散的策略中。不同于传统轨迹克隆,DP4通过动态高斯世界模型,从交互环境中学得3D空间与4D时空感知能力。该方法从单视角RGB-D观测重建当前3D场景,并预测未来3D场景,显式建模空间与时间依赖关系以优化轨迹生成。在17个仿真任务(共173种变体)和3个真实机器人任务上进行大量实验,结果表明DP4显著优于基线方法:在Adroit、DexArt和RLBench上的平均仿真任务成功率分别提升16.4%、14%和6.45%;真实机器人任务平均成功率提升8.6%。

原文摘要 · Abstract (English)

Visual imitation learning is effective for robots to learn versatile tasks. However, many existing methods rely on behavior cloning with supervised historical trajectories, limiting their 3D spatial and 4D spatiotemporal awareness. Consequently, these methods struggle to capture the 3D structures and 4D spatiotemporal relationships necessary for real-world deployment. In this work, we propose 4D Diffusion Policy (DP4), a novel visual imitation learning method that incorporates spatiotemporal awareness into diffusion-based policies. Unlike traditional approaches that rely on trajectory cloning, DP4 leverages a dynamic Gaussian world model to guide the learning of 3D spatial and 4D spatiotemporal perceptions from interactive environments. Our method constructs the current 3D scene from a single-view RGB-D observation and predicts the future 3D scene, optimizing trajectory generation by explicitly modeling both spatial and temporal dependencies. Extensive experiments across 17 simulation tasks with 173 variants and 3 real-world robotic tasks demonstrate that the 4D Diffusion Policy (DP4) outperforms baseline methods, improving the average simulation task success rate by 16.4% (Adroit), 14% (DexArt), and 6.45% (RLBench), and the average real-world robotic task success rate by 8.6%.

视觉模仿扩散模型机器人控制时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。