用单摄像头图像实现低成本机器人视觉自知,精度接近工业级。
Latent Representations for Visual Proprioception in Inexpensive Robots
- 用单帧图像通过回归模型推断机器人关节位置。
- 在6自由度低成本机器人上达到亚毫米级定位精度。
- 适合资源受限的机器人系统或快速部署场景。
机器人操作需要明确或隐含地掌握自身关节位置。高精度自知是高质量工业机器人的标配,但在非结构化环境中运行的低成本机器人中却常不可用。本文探讨:在仅有一个外部摄像头图像的简单设置下,快速、单次通过的回归架构能在多大程度上实现视觉自知?我们测试了多种潜在表示方法,包括CNN、VAE、ViT以及未标定的标记物集合,并采用适配有限数据的微调技术。通过在一台低成本6-DoF机器人上的实验,评估了可实现的精度。
原文摘要 · Abstract (English)
Robotic manipulation requires explicit or implicit knowledge of the robot's joint positions. Precise proprioception is standard in high-quality industrial robots but is often unavailable in inexpensive robots operating in unstructured environments. In this paper, we ask: to what extent can a fast, single-pass regression architecture perform visual proprioception from a single external camera image, available even in the simplest manipulation settings? We explore several latent representations, including CNNs, VAEs, ViTs, and bags of uncalibrated fiducial markers, using fine-tuning techniques adapted to the limited data available. We evaluate the achievable accuracy through experiments on an inexpensive 6-DoF robot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。