为脉冲相机设计联合压缩与感知框架,提升效率与任务性能。
A Joint Visual Compression and Perception Framework for Neuralmorphic Spiking Camera
- 分路处理空间语义与运动信息,融合生成压缩特征
- 相比最先进编码器降低17.25%码率,分类准确率提升4.3%
- 适合需要低延迟、高能效的智能视觉系统应用
神经形态脉冲相机因其极高的时间分辨率捕捉连续运动而备受关注,但其产生的二进制脉冲数据存储与传输开销巨大。针对压缩与脉冲驱动智能应用需求,我们提出脉冲智能编码(SCI)框架,对脉冲序列进行压缩并优化比特率与任务性能。受哺乳动物视觉系统启发,提出双路径架构,分别处理空间语义与运动信息,再融合生成压缩特征。引入精炼机制以保证解码特征与运动向量的一致性。进一步提出时序回归方法,整合多种运动动态,同时利用光流与形变技术。大量实验表明,该方案在脉冲压缩与分析上达到最先进水平:相比最先进编码器平均降低17.25% BD-rate,脉冲分类任务中较SpiReco提升4.3%准确率,编码端复杂度降低88.26%,推理时间节省42.41%。
原文摘要 · Abstract (English)
The advent of neuralmorphic spike cameras has garnered significant attention for their ability to capture continuous motion with unparalleled temporal resolution.However, this imaging attribute necessitates considerable resources for binary spike data storage and transmission.In light of compression and spike-driven intelligent applications, we present the notion of Spike Coding for Intelligence (SCI), wherein spike sequences are compressed and optimized for both bit-rate and task performance.Drawing inspiration from the mammalian vision system, we propose a dual-pathway architecture for separate processing of spatial semantics and motion information, which is then merged to produce features for compression.A refinement scheme is also introduced to ensure consistency between decoded features and motion vectors.We further propose a temporal regression approach that integrates various motion dynamics, capitalizing on the advancements in warping and deformation simultaneously.Comprehensive experiments demonstrate our scheme achieves state-of-the-art (SOTA) performance for spike compression and analysis.We achieve an average 17.25% BD-rate reduction compared to SOTA codecs and a 4.3% accuracy improvement over SpiReco for spike-based classification, with 88.26% complexity reduction and 42.41% inference time saving on the encoding side.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。