提升机器人抓取透明物体时的深度感知精度
ClearDepth: Enhanced Stereo Perception of Transparent Objects for Robotic Manipulation
- 基于视觉变换器与结构特征融合,增强透明物体立体深度恢复
- 在真实场景中实现高精度深度图生成,支持精准抓取
- 采用物理逼真仿真生成数据,降低真实采集成本
透明物体的深度感知在日常与物流场景中面临挑战,主要源于标准3D传感器难以准确捕捉透明或反光表面的深度信息。这一限制严重影响依赖深度图与点云的应用,尤其在机器人操作中。本文提出一种基于视觉变换器的立体深度恢复算法,并引入创新的特征后融合模块,通过图像中的结构特征提升深度恢复精度。为应对基于双目相机感知透明物体的数据集构建成本高的问题,方法结合参数对齐、领域自适应且物理真实的Sim2Real仿真,借助AI算法加速数据生成。实验表明,该模型在真实场景中表现出优异的Sim2Real泛化能力,实现了透明物体的精确深度映射,有效辅助机器人操作。项目详情见 https://sites.google.com/view/cleardepth/
原文摘要 · Abstract (English)
Transparent object depth perception poses a challenge in everyday life and logistics, primarily due to the inability of standard 3D sensors to accurately capture depth on transparent or reflective surfaces. This limitation significantly affects depth map and point cloud-reliant applications, especially in robotic manipulation. We developed a vision transformer-based algorithm for stereo depth recovery of transparent objects. This approach is complemented by an innovative feature post-fusion module, which enhances the accuracy of depth recovery by structural features in images. To address the high costs associated with dataset collection for stereo camera-based perception of transparent objects, our method incorporates a parameter-aligned, domain-adaptive, and physically realistic Sim2Real simulation for efficient data generation, accelerated by AI algorithm. Our experimental results demonstrate the model's exceptional Sim2Real generalizability in real-world scenarios, enabling precise depth mapping of transparent objects to assist in robotic manipulation. Project details are available at https://sites.google.com/view/cleardepth/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。