arXiv:2410.00368cs.CVeess.IV2024-10被引 2

为神经形态视觉系统提供带阈值的面部检测数据集,助力低功耗智能感知。

Descriptor: Face Detection Dataset for Programmable Threshold-Based Sparse-Vision

  • 基于时间差阈值生成多级事件数据,模拟神经形态传感器输出
  • 涵盖4、8、12、16四种阈值水平,支持不同灵敏度下的模型评估
  • 适配边缘计算与隐私敏感场景,推动低功耗视觉系统落地

智能焦点平面与片上图像处理已成为节能且保护隐私的嵌入式视觉系统的关键技术。然而,缺乏能真实反映神经形态传感器输出的专用数据集,制约了该技术的推广。神经形态成像器(如事件相机)可生成多种表示形式:像素地址流(记录亮度变化的时间与位置)、时间差数据、经时间差筛选/阈值化的数据、经空间变换后的图像、光流数据及统计特征等。为突破这一瓶颈,我们基于Aff-Wild2视频源,构建了一个标注完整的、基于时间阈值的面部检测数据集,提供4、8、12、16共四种阈值级别,支持对先进神经架构在不同条件下的全面评估与优化。配套工具链可从原始视频生成事件数据,提升可用性。该资源将有力推动基于时间差阈值的智能传感器视觉系统发展,实现更精准高效的目标检测与定位,促进低功耗神经形态成像技术的广泛应用。数据集已公开发布于 <a href="https://dx.doi.org/10.21227/bw2e-dj78">https://dx.doi.org/10.21227/bw2e-dj78</a>。

原文摘要 · Abstract (English)

Smart focal-plane and in-chip image processing has emerged as a crucial technology for vision-enabled embedded systems with energy efficiency and privacy. However, the lack of special datasets providing examples of the data that these neuromorphic sensors compute to convey visual information has hindered the adoption of these promising technologies. Neuromorphic imager variants, including event-based sensors, produce various representations such as streams of pixel addresses representing time and locations of intensity changes in the focal plane, temporal-difference data, data sifted/thresholded by temporal differences, image data after applying spatial transformations, optical flow data, and/or statistical representations. To address the critical barrier to entry, we provide an annotated, temporal-threshold-based vision dataset specifically designed for face detection tasks derived from the same videos used for Aff-Wild2. By offering multiple threshold levels (e.g., 4, 8, 12, and 16), this dataset allows for comprehensive evaluation and optimization of state-of-the-art neural architectures under varying conditions and settings compared to traditional methods. The accompanying tool flow for generating event data from raw videos further enhances accessibility and usability. We anticipate that this resource will significantly support the development of robust vision systems based on smart sensors that can process based on temporal-difference thresholds, enabling more accurate and efficient object detection and localization and ultimately promoting the broader adoption of low-power, neuromorphic imaging technologies. To support further research, we publicly released the dataset at \url{https://dx.doi.org/10.21227/bw2e-dj78}.

神经形态视觉事件相机低功耗面部检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。