arXiv:2511.12912cs.RO2025-11被引 2

用扩散模型模拟真实深度图噪声,实现零样本物理抓取

DiffuDepGrasp: Diffusion-based Depth Noise Modeling Empowers Sim2Real Robotic Grasping

  • 用扩散模型生成带真实传感器噪声的深度图,训练时仅需仿真数据
  • 零样本迁移下12类物体平均成功率95.7%,对未见物体也具强泛化能力
  • 部署无额外计算开销,适合对效率和鲁棒性要求高的机器人抓取场景

将基于深度的端到端策略从仿真迁移到物理机器人,可实现高效且鲁棒的抓取,但真实深度图中的空洞和噪声造成显著的模拟到现实差距,严重阻碍策略迁移。训练阶段的噪声注入或学习映射等方法因噪声模拟不真实,存在数据效率低的问题,尤其在需要精细操作或依赖成对数据集的任务中效果不佳。利用基础模型通过中间表示降低差距的方法,未能完全缓解域偏移,且部署时增加计算开销。本文解决数据效率低和部署复杂两大挑战,提出DiffuDepGrasp——一种可零样本迁移的部署高效框架,仅在仿真中训练策略。其核心创新是扩散深度生成器,通过两个协同模块合成几何纯净的仿真深度与学习到的真实传感器噪声:第一,利用时间几何先验,实现条件扩散模型的高效训练,捕捉复杂传感器噪声分布;第二,噪声嫁接模块在注入感知伪影的同时保持度量精度。部署时仅需原始深度输入,消除计算开销,在12类物体抓取任务上实现95.7%的平均成功率,并具备对未见物体的强泛化能力。

原文摘要 · Abstract (English)

Transferring the depth-based end-to-end policy trained in simulation to physical robots can yield an efficient and robust grasping policy, yet sensor artifacts in real depth maps like voids and noise establish a significant sim2real gap that critically impedes policy transfer. Training-time strategies like procedural noise injection or learned mappings suffer from data inefficiency due to unrealistic noise simulation, which is often ineffective for grasping tasks that require fine manipulation or dependency on paired datasets heavily. Furthermore, leveraging foundation models to reduce the sim2real gap via intermediate representations fails to mitigate the domain shift fully and adds computational overhead during deployment. This work confronts dual challenges of data inefficiency and deployment complexity. We propose DiffuDepGrasp, a deploy-efficient sim2real framework enabling zero-shot transfer through simulation-exclusive policy training. Its core innovation, the Diffusion Depth Generator, synthesizes geometrically pristine simulation depth with learned sensor-realistic noise via two synergistic modules. The first Diffusion Depth Module leverages temporal geometric priors to enable sample-efficient training of a conditional diffusion model that captures complex sensor noise distributions, while the second Noise Grafting Module preserves metric accuracy during perceptual artifact injection. With only raw depth inputs during deployment, DiffuDepGrasp eliminates computational overhead and achieves a 95.7% average success rate on 12-object grasping with zero-shot transfer and strong generalization to unseen objects.Project website: https://diffudepgrasp.github.io/.

机器人抓取扩散模型零样本迁移深度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。