给机器人视觉加瞄准线,提升抓取时的空间感知能力。
AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies
- 在多视角图像上叠加瞄准线和准星,直观显示机械臂位置。
- 实测提升多种策略在仿真与真实环境中的抓取成功率。
- 零成本插入式设计,不改模型架构,延迟小于1毫秒。
本文提出AimBot,一种轻量级视觉增强技术,通过在多视角RGB图像上叠加射击线和准星,为机器人抓取任务提供显式的空间引导。这些视觉线索基于深度图、相机外参和当前末端执行器位姿计算生成,明确传递了夹爪与场景中物体之间的空间关系。该方法计算开销极低(低于1毫秒),无需修改模型结构,仅需用增强后的图像替换原始图像即可。实验表明,AimBot在多种视觉运动策略中均能稳定提升性能,无论在仿真还是真实环境中,验证了空间对齐视觉反馈的有效性。
原文摘要 · Abstract (English)
In this paper, we propose AimBot, a lightweight visual augmentation technique that provides explicit spatial cues to improve visuomotor policy learning in robotic manipulation. AimBot overlays shooting lines and scope reticles onto multi-view RGB images, offering auxiliary visual guidance that encodes the end-effector's state. The overlays are computed from depth images, camera extrinsics, and the current end-effector pose, explicitly conveying spatial relationships between the gripper and objects in the scene. AimBot incurs minimal computational overhead (less than 1 ms) and requires no changes to model architectures, as it simply replaces original RGB images with augmented counterparts. Despite its simplicity, our results show that AimBot consistently improves the performance of various visuomotor policies in both simulation and real-world settings, highlighting the benefits of spatially grounded visual feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。