arXiv:2508.17230cs.CV2025-08ICCV被引 13

用4D视觉预训练提升机器人3D感知,显著提高抓取成功率

4D Visual Pre-training for Robot Learning

  • 将视觉预训练转为预测下一帧点云的扩散模型任务
  • 在12个真实场景任务中使DP3成功率平均提升28%
  • 兼容多种点云编码器,适用于大规模机器人模型

近年来,基于网络规模数据集学习的通用视觉表征在机器人操作任务中取得显著进展,实现了数据高效的机器人学习;然而这些预训练表征多基于2D图像,忽视了世界固有的3D特性。由于大规模3D数据稀缺,从网络数据集中提取通用3D表征仍具挑战。为此,我们提出一种新的4D视觉预训练框架FVP,将其视为下一帧点云预测问题,采用扩散模型建模,并直接在更大规模公开数据集上进行预训练。在十二个真实世界操作任务中,FVP使3D扩散策略(DP3)的平均成功率提升28%。经FVP预训练的DP3在模仿学习方法中达到领先水平。此外,FVP的有效性可适配多种点云编码器与数据集。最后,我们将FVP应用于RDT-1B——一个更大的视觉-语言-动作机器人模型,显著提升了其在各类机器人任务中的表现。

原文摘要 · Abstract (English)

General visual representations learned from web-scale datasets for robotics have achieved great success in recent years, enabling data-efficient robot learning on manipulation tasks; yet these pre-trained representations are mostly on 2D images, neglecting the inherent 3D nature of the world. However, due to the scarcity of large-scale 3D data, it is still hard to extract a universal 3D representation from web datasets. Instead, we are seeking a general visual pre-training framework that could improve all 3D representations as an alternative. Our framework, called FVP, is a novel 4D Visual Pre-training framework for real-world robot learning. FVP frames the visual pre-training objective as a next-point-cloud-prediction problem, models the prediction model as a diffusion model, and pre-trains the model on the larger public datasets directly. Across twelve real-world manipulation tasks, FVP boosts the average success rate of 3D Diffusion Policy (DP3) for these tasks by 28%. The FVP pre-trained DP3 achieves state-of-the-art performance across imitation learning methods. Moreover, the efficacy of FVP adapts across various point cloud encoders and datasets. Finally, we apply FVP to the RDT-1B, a larger Vision-Language-Action robotic model, enhancing its performance on various robot tasks. Our project page is available at: https://4d-visual-pretraining.github.io/

4D视觉机器人学习扩散模型点云预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。