arXiv:2606.10790cs.CV2026-06

构建首个第一人称视角下多模态手部检测数据集,提升运动场景下的检测精度。

A Multimodal RGB and Events Dataset for Hand Detection in First-Person View

  • 基于RGB数据合成事件流,模拟不同光照与尺度下的多模态数据
  • 在低光和高速运动下实现接近最先进水平的检测性能
  • 适合研究视觉-事件融合、机器人抓取与可穿戴设备方向的研究者

现有手部检测算法依赖于图像帧,受限于相机帧率。在移动机器人应用中,传统摄像头易产生运动模糊,尤其在弱光环境下。事件相机具有高动态范围、高时间分辨率和低功耗优势。近期研究表明,结合事件相机与帧相机的双目系统可提升检测准确率并优化带宽延迟平衡。然而,事件相机在目标检测任务中的主要瓶颈是训练数据不足。本文提出一种方法,基于现有的RGB Egohands数据集,利用v2e工具箱生成首个第一人称视角下的合成事件手部数据集。通过调整v2e参数,生成不同光照条件与尺度的数据版本。真实标注通过在Egohands RGB图像上微调YOLOv8模型后,再插值至高时间分辨率事件流中获得。我们使用该多模态数据集,在现有事件-视觉融合检测算法上验证,性能达到当前最优水平。

原文摘要 · Abstract (English)

Existing hand detection algorithms work on images and the detection rate is restricted by the frame rate of the camera. In hand detection applications for moving robotic systems, conventional cameras cause motion blur, especially in darker lighting conditions. We can leverage the use of event-based cameras which possess a high dynamic range, high temporal resolution, and low power consumption. Recent work has shown that using a stereo setup of an event-based and a frame-based camera improves detection accuracy and the bandwidth-latency tradeoff. The main bottleneck in using event-based cameras in object detection and recognition tasks is a relatively low amount of training data. In this work, we propose a methodology and an exemplary synthetic event-based hand dataset from an egocentric, first-person view perspective. The data is synthesized from the existing RGB Egohands dataset with the v2e toolbox. Parameters of the v2e toolbox are varied to provide versions of the dataset with different lighting conditions and scales. Ground truth detections are generated with a fine-tuned YOLOv8 model which is applied to the RGB images in the Egohands dataset and interpolated on the high-temporal resolution events. We use the multi-modal dataset to perform hand detection with existing object detection algorithms which use a multi-modal setup of event and RGB cameras and demonstrate performance comparable to the state-of-the-art.

手部检测事件相机多模态第一人称

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。