用空间感知+扩散模型提升机器人抓取的泛化能力
Spatial RoboGrasp: Generalized Robotic Grasping Control Policy
- 融合深度估计与6-DoF抓取提示,构建统一空间表征
- 在环境变化下抓取成功率提升40%,任务成功率提高45%
- 适合需要高鲁棒性的工业抓取与复杂场景应用
实现跨多样化环境的通用且精确的机器人操作仍是重大挑战,主要受限于空间感知能力。现有基于模仿学习的方法多依赖原始RGB输入和手工特征,易过拟合,且在光照、遮挡和物体条件变化下3D推理能力差。本文提出统一框架,结合鲁棒多模态感知与可靠抓取预测:通过领域随机化增强、单目深度估计及深度感知的6-DoF抓取提示,生成统一空间表示用于下游动作规划。该表征结合高层任务提示,驱动基于扩散模型的策略,生成精准动作序列,在环境变化下抓取成功率提升40%,任务成功率提高45%。结果表明,基于空间定位感知与扩散式模仿学习的方案,为通用机器人抓取提供了可扩展且鲁棒的解决方案。
原文摘要 · Abstract (English)
Achieving generalizable and precise robotic manipulation across diverse environments remains a critical challenge, largely due to limitations in spatial perception. While prior imitation-learning approaches have made progress, their reliance on raw RGB inputs and handcrafted features often leads to overfitting and poor 3D reasoning under varied lighting, occlusion, and object conditions. In this paper, we propose a unified framework that couples robust multimodal perception with reliable grasp prediction. Our architecture fuses domain-randomized augmentation, monocular depth estimation, and a depth-aware 6-DoF Grasp Prompt into a single spatial representation for downstream action planning. Conditioned on this encoding and a high-level task prompt, our diffusion-based policy yields precise action sequences, achieving up to 40% improvement in grasp success and 45% higher task success rates under environmental variation. These results demonstrate that spatially grounded perception, paired with diffusion-based imitation learning, offers a scalable and robust solution for general-purpose robotic grasping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。