arXiv:2503.04862cs.CVcs.RO2025-03被引 1

用Transformer实现人形机器人对微小物体的高精度视觉定位

High-Precision Transformer-Based Visual Servoing for Humanoid Robots in Aligning Tiny Objects

  • 融合头部和躯干摄像头与关节角度,基于Transformer进行视觉伺服控制
  • 对M4-M8螺丝定位平均误差0.8-1.3毫米,成功率93%-100%
  • 适合需要精密操作的人形机器人任务,如装配与维修

真实世界中,人形机器人对微小物体的高精度对齐仍是一个普遍且关键的挑战。本文提出一种基于视觉的框架,用于精确估计并控制手持工具与目标物体之间的相对位置,例如螺丝刀头与螺丝槽。通过融合机器人头部和躯干摄像头的图像及其头部关节角度,所提出的基于Transformer的视觉伺服方法能有效校正近距离下的位置误差。在M4-M8螺丝上的实验表明,平均收敛误差为0.8-1.3毫米,成功率高达93%-100%。对比分析验证了该高精度对齐能力得益于本文提出的距离估计Transformer架构和多感知头机制。

原文摘要 · Abstract (English)

High-precision tiny object alignment remains a common and critical challenge for humanoid robots in real-world. To address this problem, this paper proposes a vision-based framework for precisely estimating and controlling the relative position between a handheld tool and a target object for humanoid robots, e.g., a screwdriver tip and a screw head slot. By fusing images from the head and torso cameras on a robot with its head joint angles, the proposed Transformer-based visual servoing method can correct the handheld tool's positional errors effectively, especially at a close distance. Experiments on M4-M8 screws demonstrate an average convergence error of 0.8-1.3 mm and a success rate of 93\%-100\%. Through comparative analysis, the results validate that this capability of high-precision tiny object alignment is enabled by the Distance Estimation Transformer architecture and the Multi-Perception-Head mechanism proposed in this paper.

视觉伺服人形机器人Transformer精密对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。