用无标注真实数据提升机器人对箱子姿态与形状的感知能力
Box Pose and Shape Estimation and Domain Adaptation for Large-Scale Warehouse Automation
- 通过自监督域适应,利用真实未标注数据优化感知模型
- 在5万张真实图像上测试,显著优于纯仿真训练模型
- 适合大规模仓储自动化中需低成本标注的场景
现代仓储自动化系统依赖大量智能机器人生成海量数据,其中多数未被标注。本文提出一种自监督域适应流程,利用真实世界中的无标签数据提升感知模型性能,无需人工标注。研究聚焦于箱子姿态与形状估计,构建了‘正确并验证’的自监督盒子姿态与形状估计框架。我们在多种模拟和真实工业环境中进行广泛评估,包括一个包含5万张图像的大规模真实数据集。自监督模型显著优于仅在仿真中训练的模型,并大幅超越零样本3D边界框估计基线。
原文摘要 · Abstract (English)
Modern warehouse automation systems rely on fleets of intelligent robots that generate vast amounts of data -- most of which remains unannotated. This paper develops a self-supervised domain adaptation pipeline that leverages real-world, unlabeled data to improve perception models without requiring manual annotations. Our work focuses specifically on estimating the pose and shape of boxes and presents a correct-and-certify pipeline for self-supervised box pose and shape estimation. We extensively evaluate our approach across a range of simulated and real industrial settings, including adaptation to a large-scale real-world dataset of 50,000 images. The self-supervised model significantly outperforms models trained solely in simulation and shows substantial improvements over a zero-shot 3D bounding box estimation baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。