首个融合视觉与热成像的脉冲动作识别数据集,助力低功耗视频理解研究。
SPACT18: Spiking Human Action Recognition Benchmark Dataset with Complementary RGB and Thermal Modalities
- 构建首个使用脉冲相机的多模态动作识别数据集,含同步RGB与热成像。
- 数据保留脉冲信号稀疏性与高时间精度,支持神经网络直接建模。
- 适合关注低功耗视觉、脉冲神经网络及多模态融合的研究者。
脉冲相机是一种类生物视觉传感器,通过像素级光强累积异步触发脉冲,具备超高的能效和出色的时序分辨率。与记录光强变化的事件相机不同,脉冲相机能提供更精细的时空分辨率和对连续变化的更精确表征。本文首次提出一个基于脉冲相机的视频动作识别(VAR)数据集,并同步包含RGB与热成像模态,为脉冲神经网络(SNNs)提供全面的基准测试平台。通过保持脉冲数据的固有稀疏性和时间精度,本研究构建的三个数据集为探索多模态视频理解提供了独特平台,并可直接用于比较脉冲、热成像与RGB模态的表现。该工作贡献了一个新数据集,将推动面向动作识别任务的节能、超低功耗视频理解研究。
原文摘要 · Abstract (English)
Spike cameras, bio-inspired vision sensors, asynchronously fire spikes by accumulating light intensities at each pixel, offering ultra-high energy efficiency and exceptional temporal resolution. Unlike event cameras, which record changes in light intensity to capture motion, spike cameras provide even finer spatiotemporal resolution and a more precise representation of continuous changes. In this paper, we introduce the first video action recognition (VAR) dataset using spike camera, alongside synchronized RGB and thermal modalities, to enable comprehensive benchmarking for Spiking Neural Networks (SNNs). By preserving the inherent sparsity and temporal precision of spiking data, our three datasets offer a unique platform for exploring multimodal video understanding and serve as a valuable resource for directly comparing spiking, thermal, and RGB modalities. This work contributes a novel dataset that will drive research in energy-efficient, ultra-low-power video understanding, specifically for action recognition tasks using spike-based data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。