arXiv:2410.10592cs.AReess.IV2024-10被引 1

无需模数转换的像素内计算架构,提升边缘视觉能效与带宽

Voltage-Controlled Magnetic Tunnel Junction based ADC-less Global Shutter Processing-in-Pixel for Extreme-Edge Intelligence

  • 在像素级直接计算,用电压控磁隧道结实现无损读取
  • 相较传统系统,前端能耗降8.2倍,通信能耗降8.5倍,带宽减6倍
  • 适合高能效边缘智能场景,如物联网视觉感知

摄像头传感器产生的海量数据推动了边缘设备上计算机视觉任务的节能处理方案研究。处理-在像素(processing-in-pixel)通过在像素阵列内集成大规模并行模拟计算能力,将第一层神经网络输出激活值而非原始传感数据作为输出,显著提升能效和带宽效率。本文提出一种无模数转换(ADC-less)的高效处理-在像素架构,采用基于霍耶正则化训练的优化二值激活神经网络,在复杂视觉任务中保持高精度。同时,提出一种利用纳米级电压控磁隧道结(VC-MTJ)实现快速且无干扰读取的全局快门突发存储读取方案。此外,构建了融合器件与电路约束(如器件开关特性、电路非线性)的算法框架,基于先进制造工艺(GlobalFoundries 22nm FDX)的实测VC-MTJ特性与大量电路仿真。最终在CIFAR10和ImageNet两个复杂数据集上评估,系统前端与通信能耗分别降低8.2倍和8.5倍,带宽减少6倍,测试准确率无明显下降。

原文摘要 · Abstract (English)

The vast amount of data generated by camera sensors has prompted the exploration of energy-efficient processing solutions for deploying computer vision tasks on edge devices. Among the various approaches studied, processing-in-pixel integrates massively parallel analog computational capabilities at the extreme-edge, i.e., within the pixel array and exhibits enhanced energy and bandwidth efficiency by generating the output activations of the first neural network layer rather than the raw sensory data. In this article, we propose an energy and bandwidth efficient ADC-less processing-in-pixel architecture. This architecture implements an optimized binary activation neural network trained using Hoyer regularizer for high accuracy on complex vision tasks. In addition, we also introduce a global shutter burst memory read scheme utilizing fast and disturb-free read operation leveraging innovative use of nanoscale voltage-controlled magnetic tunnel junctions (VC-MTJs). Moreover, we develop an algorithmic framework incorporating device and circuit constraints (characteristic device switching behavior and circuit non-linearity) based on state-of-the-art fabricated VC-MTJ characteristics and extensive circuit simulations using commercial GlobalFoundries 22nm FDX technology. Finally, we evaluate the proposed system's performance on two complex datasets - CIFAR10 and ImageNet, showing improvements in front-end and communication energy efficiency by 8.2x and 8.5x respectively and reduction in bandwidth by 6x compared to traditional computer vision systems, without any significant drop in the test accuracy.

边缘智能像素计算磁隧道结能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。