首次对比商用微神经处理器,揭示硬件规格与实测性能的惊人差异
Benchmarking Ultra-Low-Power $μ$NPUs
- 构建统一编译管道,实现量化模型跨平台公平测试
- 多款μNPU实测性能与规格不符,部分出现反常复杂度响应
- 为硬件与软件开发者提供实用参考,助力低功耗边缘计算
高效本地化神经网络推理可带来可预测延迟、更好隐私保护和更低运营成本,优于云端推理。这推动了微控制器级神经网络加速器(即μNPUs)的发展,专为超低功耗应用设计。本文首次对多个商用μNPUs进行系统性对比评估,包含多个平台的独立基准测试。为确保公平性,我们开发并开源了一套模型编译管道,支持在不同微控制器硬件上对量化模型进行一致基准测试。分析结果揭示了预期性能趋势,也发现了硬件规格与实际表现之间的意外差异,包括某些μNPUs在模型复杂度增加时表现出异常的扩展行为。该工作为μNPU平台的持续评估奠定了基础,并为软硬件开发者提供了实用洞见。
原文摘要 · Abstract (English)
Efficient on-device neural network (NN) inference offers predictable latency, improved privacy and reliability, and lower operating costs for vendors than cloud-based inference. This has sparked recent development of microcontroller-scale NN accelerators, also known as neural processing units ($μ$NPUs), designed specifically for ultra-low-power applications. We present the first comparative evaluation of a number of commercially-available $μ$NPUs, including the first independent benchmarks for multiple platforms. To ensure fairness, we develop and open-source a model compilation pipeline supporting consistent benchmarking of quantized models across diverse microcontroller hardware. Our resulting analysis uncovers both expected performance trends as well as surprising disparities between hardware specifications and actual performance, including certain $μ$NPUs exhibiting unexpected scaling behaviors with model complexity. This work provides a foundation for ongoing evaluation of $μ$NPU platforms, alongside offering practical insights for both hardware and software developers in this rapidly evolving space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。