arXiv:2409.19389cs.DCcs.AI2024-09

NV-1芯片实现超低功耗下10倍以上并行加速,突破边缘设备算力瓶颈。

Co-design of a novel CMOS highly parallel, low-power, multi-chip neural network accelerator

  • 采用非冯诺依曼架构,集成多组处理器-内存单元实现高度并行计算。
  • 实测功耗降低100倍以上,处理速度提升10倍以上,无传统内存瓶颈。
  • 适合部署在无外接电源的智能传感器等边缘设备,如安防摄像头、语音助手。

为何安防摄像头、传感器和Siri仍依赖云端计算?核心限制在于缺乏超低功耗、高性能的专用芯片。本文提出NV-1,一款新型低功耗ASIC AI处理器,通过大量并行的处理器-内存单元组合(即大幅非冯诺依曼架构),实现超过10倍的并行处理加速与超过100倍的能耗降低,避免了传统单体内存带来的性能瓶颈。该原型芯片由算法、软件与电路设计团队协同开发,创新通信协议消除地址总线,显著降低节点间数据传输开销。早期构建数字孪生模型确保软硬件对齐,实测性能与预测一致。目前该芯片已用于实际边缘传感器场景,多项验证正在推进,证明其在真实世界中具备超低功耗、高性能的可行性。

原文摘要 · Abstract (English)

Why do security cameras, sensors, and siri use cloud servers instead of on-board computation? The lack of very-low-power, high-performance chips greatly limits the ability to field untethered edge devices. We present the NV-1, a new low-power ASIC AI processor that greatly accelerates parallel processing (> 10X) with dramatic reduction in energy consumption (> 100X), via many parallel combined processor-memory units, i.e., a drastically non-von-Neumann architecture, allowing very large numbers of independent processing streams without bottlenecks due to typical monolithic memory. The current initial prototype fab arises from a successful co-development effort between algorithm- and software-driven architectural design and VLSI design realities. An innovative communication protocol minimizes power usage, and data transport costs among nodes were vastly reduced by eliminating the address bus, through local target address matching. Throughout the development process, the software and architecture teams were able to innovate alongside the circuit design team's implementation effort. A digital twin of the proposed hardware was developed early on to ensure that the technical implementation met the architectural specifications, and indeed the predicted performance metrics have now been thoroughly verified in real hardware test data. The resulting device is currently being used in a fielded edge sensor application; additional proofs of principle are in progress demonstrating the proof on the ground of this new real-world extremely low-power high-performance ASIC device.

边缘计算低功耗AI芯片非冯诺依曼

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。