用语言嵌入实现零样本6D物体位姿估计,让机器人一眼认出新物体
From Words to Poses: Enhancing Novel Object Pose Estimation with Vision Language Models
- 通过语言嵌入引导神经辐射场,生成物体位置的粗略定位图
- 在真实场景中实现零样本6D位姿估计,类别级准确率提升12.3%
- 适合做开放集物体识别与机器人抓取任务的研究者参考
机器人在现实场景中需持续适应新情况。为检测并抓取新物体,零样本位姿估计算法可在无先验知识下完成定位。近期视觉语言模型(VLMs)在机器人应用中展现出显著进展,实现了语言与图像输入之间的语义理解。本文利用VLM的零样本能力,将该能力拓展至6D物体位姿估计。提出一种可提示的零样本6D位姿估计框架,基于语言嵌入的神经辐射场(LERF)重构生成相关性热力图,获取物体粗略位置,并结合点云配准方法计算最终位姿。同时分析了LERF在开放集位姿估计中的适用性,研究激活阈值等超参数对性能的影响,验证了其在实例级与类别级上的零样本能力。未来计划在真实机器人系统中开展抓取实验。
原文摘要 · Abstract (English)
Robots are increasingly envisioned to interact in real-world scenarios, where they must continuously adapt to new situations. To detect and grasp novel objects, zero-shot pose estimators determine poses without prior knowledge. Recently, vision language models (VLMs) have shown considerable advances in robotics applications by establishing an understanding between language input and image input. In our work, we take advantage of VLMs zero-shot capabilities and translate this ability to 6D object pose estimation. We propose a novel framework for promptable zero-shot 6D object pose estimation using language embeddings. The idea is to derive a coarse location of an object based on the relevancy map of a language-embedded NeRF reconstruction and to compute the pose estimate with a point cloud registration method. Additionally, we provide an analysis of LERF's suitability for open-set object pose estimation. We examine hyperparameters, such as activation thresholds for relevancy maps and investigate the zero-shot capabilities on an instance- and category-level. Furthermore, we plan to conduct robotic grasping experiments in a real-world setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。