arXiv:2506.08974eess.IV2025-06被引 7

公开3万张水下手势图像数据集,助力人机水下交互研究

Diver-Robot Communication Dataset for Underwater Hand Gesture Recognition

  • 融合视觉与手套传感双模态数据,支持水下手势识别
  • 包含近900个手势、5名潜水员在海池环境下的多距离采集数据
  • 适合研究水下视觉识别、人机协同及低可见度通信系统

本文发布一个用于水下人机交互的手势图像数据集。通过开放获取该数据集,旨在探索自主水下航行器(AUV)通过视觉检测潜水员手势作为通信方式的可行性。除图像记录外,同一数据集还通过智能手势识别手套同步采集,该手套采用弹性体传感器和本地处理,通过声学方式将手势指令传给AUV。尽管此方法可在不同能见度条件下使用甚至无视线时运行,但存在声学传输带来的通信延迟。为对比效率,手套内置了名为CADDIAN的手势语言视觉标记,并与水下摄像头同步记录。数据集包含超过3万帧图像,涵盖近900个手势,按片段分类标注。数据在海、池环境中由5名潜水员分别以1、2、3米距离采集,保持平衡分布。报告了手套识别的平均反应时间、手势执行时间、识别成功率、传输时间等统计结果。该数据集可为不同能见度条件下先进视觉手势识别技术的性能比较提供基准。

原文摘要 · Abstract (English)

In this paper, we present a dataset of diving gesture images used for human-robot interaction underwater. By offering this open access dataset, the paper aims at investigating the potential of using visual detection of diving gestures from an autonomous underwater vehicle (AUV) as a form of communication with a human diver. In addition to the image recording, the same dataset was recorded using a smart gesture recognition glove. The glove uses elastomer sensors and on-board processing to determine the selected gesture and transmit the command associated with the gesture to the AUV via acoustics. Although this method can be used under different visibility conditions and even without line of sight, it introduces a communication delay required for the acoustic transmission of the gesture command. To compare efficiency, the glove was equipped with visual markers proposed in a gesture-based language called CADDIAN and recorded with an underwater camera in parallel to the glove's onboard recognition process. The dataset contains over 30,000 underwater frames of nearly 900 individual gestures annotated in corresponding snippet folders. The dataset was recorded in a balanced ratio with five different divers in sea and five different divers in pool conditions, with gestures recorded at 1, 2 and 3 metres from the camera. The glove gesture recognition statistics are reported in terms of average diver reaction time, average time taken to perform a gesture, recognition success rate, transmission times and more. The dataset presented should provide a good baseline for comparing the performance of state of the art visual diving gesture recognition techniques under different visibility conditions.

水下通信手势识别人机交互数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。