提出自适应时间衰减表面,让事件相机在高速下更准地识别人脸和关键点。
Locally Adaptive Decay Surfaces for High-Speed Face and Landmark Detection with Event Cameras
- 按局部信号动态调节时间衰减,避免全局参数导致的模糊或失真。
- 在30Hz和240Hz下均提升人脸检测与关键点精度,240Hz时误差低于1%。
- 适合需要实时高帧率的人机交互系统,尤其擅长处理快速运动场景。
事件相机以微秒级分辨率记录亮度变化,但将其稀疏异步输出转化为神经网络可利用的密集张量仍是核心挑战。传统直方图或全局衰减时间表面采用固定时间参数,导致静止期结构保留与快速运动边缘清晰度之间存在权衡。本文提出局部自适应衰减表面(LADS),在每个位置根据局部信号动态调节时间衰减。探索了基于事件速率、高斯拉普拉斯响应和高频谱能量三种策略,可在静止区域保留细节,同时减少密集活动区域的模糊。在公开数据集上大量实验表明,相比标准非自适应表示,LADS持续提升人脸检测与面部关键点准确率。在30Hz下,检测准确率更高且关键点误差更低;在240Hz时,缓解了高频率下通常出现的性能下降,关键点归一化平均误差维持在2.44%,人脸检测mAP50达0.966。这些高频结果甚至超越此前在30Hz下运行的工作,为事件相机人脸分析树立新基准。此外,由于在表示阶段保留空间结构,LADS支持使用更轻量网络架构,仍保持实时性能。结果凸显上下文感知时间融合在类脑视觉中的重要性,指向利用事件相机优势实现实时高帧率人机交互系统。
原文摘要 · Abstract (English)
Event cameras record luminance changes with microsecond resolution, but converting their sparse, asynchronous output into dense tensors that neural networks can exploit remains a core challenge. Conventional histograms or globally-decayed time-surface representations apply fixed temporal parameters across the entire image plane, which in practice creates a trade-off between preserving spatial structure during still periods and retaining sharp edges during rapid motion. We introduce Locally Adaptive Decay Surfaces (LADS), a family of event representations in which the temporal decay at each location is modulated according to local signal dynamics. Three strategies are explored, based on event rate, Laplacian-of-Gaussian response, and high-frequency spectral energy. These adaptive schemes preserve detail in quiescent regions while reducing blur in regions of dense activity. Extensive experiments on the public data show that LADS consistently improves both face detection and facial landmark accuracy compared to standard non-adaptive representations. At 30 Hz, LADS achieves higher detection accuracy and lower landmark error than either baseline, and at 240 Hz it mitigates the accuracy decline typically observed at higher frequencies, sustaining 2.44 % normalized mean error for landmarks and 0.966 mAP50 in face detection. These high-frequency results even surpass the accuracy reported in prior works operating at 30 Hz, setting new benchmarks for event-based face analysis. Moreover, by preserving spatial structure at the representation stage, LADS supports the use of much lighter network architectures while still retaining real-time performance. These results highlight the importance of context-aware temporal integration for neuromorphic vision and point toward real-time, high-frequency human-computer interaction systems that exploit the unique advantages of event cameras.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。