arXiv:2603.08725cs.ARcs.CV2026-03中稿 · the IEEE Internati…综述被引 9

对比三类低功耗芯片,揭示传感器端计算的能效优势。

Performance Analysis of Edge and In-Sensor AI Processors: A Comparative Review

  • 按计算范式分类主流边缘与传感处理器,分析其适用场景。
  • 在3.36亿次运算模型上实测,传感器芯片能效比最高,延迟最低。
  • 适合关注低功耗智能硬件、边缘AI部署的研究者与工程师。

本文综述了超低功耗边缘处理器的快速发展,涵盖异构片上系统(SoCs)、神经加速器、近传感器与传感器内架构,以及新兴的数据流与内存中心设计。根据计算范式、功耗范围和存储层级,对商用与研究级平台进行分类,并分析其在持续运行与低延迟人工智能任务中的适用性。为补充架构分析,我们在三种代表性处理器上测试了3.36亿次乘加运算(MAC)的分割模型(PicoSAM2):GAP9采用多核RISC-V架构并配备硬件加速器;STM32N6结合先进的ARM Cortex-M55核心与专用神经架构加速器;索尼IMX500代表传感器内堆叠互补金属氧化物半导体(CMOS)计算。三者分别覆盖微控制器级、嵌入式神经加速器与传感器内计算范式。评估指标包括延迟、推理效率、能量效率与能量-延迟积。结果表明硬件行为显著分化:IMX500实现86.2 MAC/周期利用率与最低能量-延迟积,凸显传感器内处理的技术成熟度;GAP9在微控制器级功耗预算下能效最优;而STM32N6虽能耗更高,但提供最低原始延迟。综述与实测共同揭示当前超低功耗与传感器内AI处理器的设计趋势与实用权衡。

原文摘要 · Abstract (English)

This review examines the rapidly evolving landscape of ultra-low-power edge processors, covering heterogeneous Systems-on-Chips (SoCs), neural accelerators, near-sensor and in-sensor architectures, and emerging dataflow and memory-centric designs. We categorize commercially available and research-grade platforms according to their compute paradigms, power envelopes, and memory hierarchies, and analyze their suitability for always-on and latency-critical Artificial Intelligence (AI) workloads. To complement the architectural overview with empirical evidence, we benchmark a 336 million Multiply-Accumulate (MAC) segmentation model (PicoSAM2) on three representative processors: GAP9, leveraging a multi-core RISC-V architecture augmented with hardware accelerators; the STM32N6, which pairs an advanced ARM Cortex-M55 core with a dedicated neural architecture accelerator; and the Sony IMX500, representing in-sensor stacked-Complementary Metal-Oxide-Semiconductor (CMOS) compute. Collectively, these platforms span MCU-class, embedded neural accelerator, and in-sensor paradigms. The evaluation reports latency, inference efficiency, energy efficiency, and energy-delay product. The results show a clear divergence in hardware behavior, with the IMX500 achieving the highest utilization (86.2 MAC/cycle) and the lowest energy-delay product, highlighting the growing significance and technological maturity of in-sensor processing. GAP9 offers the best energy efficiency within microcontroller-class power budgets, and the STM32N6 provides the lowest raw latency at a significantly higher energy cost. Together, the review and benchmarks provide a unified view of the current design directions and practical trade-offs that are shaping the next generation of ultra-low-power and in-sensor AI processors.

边缘计算传感器计算低功耗AI芯片

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。