用神经种群编码实现快速精准的物体位姿估计
Object-Pose Estimation With Neural Population Codes
- 用神经种群编码直接映射感官输入到物体旋转
- 在T-LESS数据集上达84.7%精度,推理仅需3.2毫秒
- 适合需要高速高精度位姿估计的机器人装配任务
机器人装配任务需要精确的物体位姿估计,尤其当避免使用昂贵的机械约束时。物体对称性使得感官输入到旋转的直接映射变得模糊且无唯一训练目标。现有方法如评估多个位姿假设或预测概率分布,存在显著计算开销。本文表明,采用神经种群编码表示物体旋转可克服这些限制,实现旋转的直接映射与端到端学习。结果表明,该方法支持快速准确的位姿估计:在T-LESS数据集上,仅使用灰度图像输入,苹果M1 CPU上推理时间仅为3.2毫秒,最大对称感知表面距离精度达84.7%,优于直接映射到位姿的69.7%。
原文摘要 · Abstract (English)
Robotic assembly tasks require object-pose estimation, particularly for tasks that avoid costly mechanical constraints. Object symmetry complicates the direct mapping of sensory input to object rotation, as the rotation becomes ambiguous and lacks a unique training target. Some proposed solutions involve evaluating multiple pose hypotheses against the input or predicting a probability distribution, but these approaches suffer from significant computational overhead. Here, we show that representing object rotation with a neural population code overcomes these limitations, enabling a direct mapping to rotation and end-to-end learning. As a result, population codes facilitate fast and accurate pose estimation. On the T-LESS dataset, we achieve inference in 3.2 milliseconds on an Apple M1 CPU and a Maximum Symmetry-Aware Surface Distance accuracy of 84.7% using only gray-scale image input, compared to 69.7% accuracy when directly mapping to pose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。