用低成本相机自动识别金属物体,减少人工标注
Reasoning and Learning a Perceptual Metric for Self-Training of Reflective Objects in Bin-Picking with a Low-cost Camera
- 通过推理约束优化物体姿态猜测
- 自训练使模型能处理未见金属物
- 适合工业自动化场景应用
使用低成本RGB-D相机进行金属物体抓取时,常因深度信息稀疏和反光表面导致误检,需人工标注。为此提出两阶段框架:第一阶段采用多物体姿态推理(MoPR)算法,在深度、碰撞和边界约束下优化姿态假设;第二阶段引入对称性感知的李群贝叶斯高斯混合模型(SaL-BGMM)结合期望最大化(EM)算法,实现对姿态候选的对称性过滤。此外,设计加权排名信息噪声对比损失(WR-InfoNCE),使低成本相机从重构数据中学习感知度量,支持对未训练或未见过物体的自训练。实验表明,该方法在ROBI数据集及新构建的Self-ROBI数据集上均优于多个前沿方法。
原文摘要 · Abstract (English)
Bin-picking of metal objects using low-cost RGB-D cameras often suffers from sparse depth information and reflective surface textures, leading to errors and the need for manual labeling. To reduce human intervention, we propose a two-stage framework consisting of a metric learning stage and a self-training stage. Specifically, to automatically process data captured by a low-cost camera (LC), we introduce a Multi-object Pose Reasoning (MoPR) algorithm that optimizes pose hypotheses under depth, collision, and boundary constraints. To further refine pose candidates, we adopt a Symmetry-aware Lie-group based Bayesian Gaussian Mixture Model (SaL-BGMM), integrated with the Expectation-Maximization (EM) algorithm, for symmetry-aware filtering. Additionally, we propose a Weighted Ranking Information Noise Contrastive Estimation (WR-InfoNCE) loss to enable the LC to learn a perceptual metric from reconstructed data, supporting self-training on untrained or even unseen objects. Experimental results show that our approach outperforms several state-of-the-art methods on both the ROBI dataset and our newly introduced Self-ROBI dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。