arXiv:2503.04308cs.ROcs.CV2025-03中稿 · Presentation at th…被引 1

构建真实场景玻璃数据集,提升机器人识别透明杯能力

Shaken, Not Stirred: A Novel Dataset for Visual Understanding of Glasses in Human-Robot Bartending Tasks

  • 用深度信息自动标注多视角RGB-D图像,减少人工成本
  • 新数据集含7850张图,训练模型在机器人调酒任务中达81%成功率
  • 解决透明/反光玻璃识别难题,适合具身智能与机器人视觉研究

现有物体检测数据集对玻璃类物体的多样性覆盖不足,因其透明和反光特性导致检测困难。尤其开放词汇检测器在区分玻璃子类别时表现不佳,影响机器人在感知、规划与执行中的准确性。本文提出一种基于RGB-D传感器的低成本真实世界数据采集方法,设计自动标注流程,利用深度测量生成所有帧的标签。构建了名为GlassNICOL Dataset的新数据集,由神经启发协作平台NICOL采集,包含来自五个摄像头的7850张图像。实验表明,所训练基线模型优于现有先进开放词汇方法。进一步将该模型部署于实体机器人系统,在人机调酒任务中实现81%的成功率。

原文摘要 · Abstract (English)

Datasets for object detection often do not account for enough variety of glasses, due to their transparent and reflective properties. Specifically, open-vocabulary object detectors, widely used in embodied robotic agents, fail to distinguish subclasses of glasses. This scientific gap poses an issue for robotic applications that suffer from accumulating errors between detection, planning, and action execution. This paper introduces a novel method for acquiring real-world data from RGB-D sensors that minimizes human effort. We propose an auto-labeling pipeline that generates labels for all the acquired frames based on the depth measurements. We provide a novel real-world glass object dataset GlassNICOLDataset that was collected on the Neuro-Inspired COLlaborator (NICOL), a humanoid robot platform. The dataset consists of 7850 images recorded from five different cameras. We show that our trained baseline model outperforms state-of-the-art open-vocabulary approaches. In addition, we deploy our baseline model in an embodied agent approach to the NICOL platform, on which it achieves a success rate of 81% in a human-robot bartending scenario.

机器人视觉玻璃识别数据集具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。