用投影光栅实现硬盘自动拆解的高精度3D视觉感知
Fringe Projection Based Vision Pipeline for Autonomous Hard Drive Disassembly

- 用光栅投影获取3D深度,失败时自动触发补全模块
- 实例分割准确率96%,深度误差仅2.317毫米
- 实时运行,适合工业机器人拆解场景
未回收的电子垃圾造成重大经济损失。硬盘驱动器(HDD)是高价值电子垃圾,需机器人自动化拆解。现有方法在3D感知、场景理解与紧固件定位方面存在碎片化、鲁棒性差等问题。本文提出一种自主视觉流水线:采用光栅投影轮廓术(FPP)进行3D感知,当FPP失效时选择性触发深度补全模块,并集成轻量级实时实例分割网络实现场景理解与关键部件定位。利用同一套FPP相机-投影仪系统完成深度感知与组件定位,使深度图与分割掩码像素级对齐,无需注册,优于常见RGB-D系统。优化深度补全与分割网络以适配部署推理。系统在实例分割上取得box mAP@50 0.960、mask mAP@50 0.957;深度补全配置(基于Depth Anything V2 Base)RMSE为2.317 mm,MAE为1.836 mm;Platter Facing推理栈在测试机上实现12.86毫秒延迟与77.7帧每秒吞吐率。采用模拟到真实迁移学习扩充物理数据集。该感知流水线提供高保真语义与空间信息,适用于下游机器人拆解。用于HDD实例分割的合成数据集将公开共享。
原文摘要 · Abstract (English)
Unrecovered e-waste represents a significant economic loss. Hard disk drives (HDDs) comprise a valuable e-waste stream necessitating robotic disassembly. Automating the disassembly of HDDs requires holistic 3D sensing, scene understanding, and fastener localization, however current methods are fragmented, lack robust 3D sensing, and lack fastener localization. We propose an autonomous vision pipeline which performs 3D sensing using a Fringe Projection Profilometry (FPP) module, with selective triggering of a depth completion module where FPP fails, and integrates this module with a lightweight, real-time instance segmentation network for scene understanding and critical component localization. By utilizing the same FPP camera-projector system for both our depth sensing and component localization modules, our depth maps and derived 3D geometry are inherently pixel-wise aligned with the segmentation masks without registration, providing an advantage over RGB-D perception systems common in industrial sensing. We optimize both our trained depth completion and instance segmentation networks for deployment-oriented inference. The proposed system achieves a box mAP@50 of 0.960 and mask mAP@50 of 0.957 for instance segmentation, while the selected depth completion configuration with the Depth Anything V2 Base backbone achieves an RMSE of 2.317 mm and MAE of 1.836 mm; the Platter Facing learned inference stack achieved a combined latency of 12.86 ms and a throughput of 77.7 Frames Per Second (FPS) on the evaluation workstation. Finally, we adopt a sim-to-real transfer learning approach to augment our physical dataset. The proposed perception pipeline provides both high-fidelity semantic and spatial data which can be valuable for downstream robotic disassembly. The synthetic dataset developed for HDD instance segmentation will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。