eIQ Neutron通过软硬协同设计,让边缘AI推理速度提升3.3倍。
eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations
- 采用数据驱动架构与约束编程编译器,动态优化计算与数据流动。
- 在同等算力和内存下,平均提速1.8倍,峰值达4倍。
- 适合对能效和推理速度敏感的嵌入式AI场景。
神经处理器(NPUs)是实现资源受限边缘环境高效AI推理的关键。尽管峰值千兆操作每秒(TOPS)常被用作性能指标,但其无法准确反映真实表现,且往往与更高芯片成本相关。为此,架构设计需聚焦于最大化计算利用率,同时保持灵活性。本文提出集成于商用旗舰MPU中的eIQ Neutron高效NPU,以及协同设计的编译算法。该架构采用灵活的数据驱动设计,编译器则通过约束编程方法,根据工作负载特性优化计算与数据移动。相比领先的嵌入式NPU及编译器栈,在标准AI基准测试中,本方案在相同TOPS与内存资源下实现平均1.8倍加速(峰值达4倍)。即使面对算力与内存均翻倍的NPU,Neutron仍可提供最高3.3倍的性能优势。
原文摘要 · Abstract (English)
Neural Processing Units (NPUs) are key to enabling efficient AI inference in resource-constrained edge environments. While peak tera operations per second (TOPS) is often used to gauge performance, it poorly reflects real-world performance and typically rather correlates with higher silicon cost. To address this, architects must focus on maximizing compute utilization, without sacrificing flexibility. This paper presents the eIQ Neutron efficient-NPU, integrated into a commercial flagship MPU, alongside co-designed compiler algorithms. The architecture employs a flexible, data-driven design, while the compiler uses a constrained programming approach to optimize compute and data movement based on workload characteristics. Compared to the leading embedded NPU and compiler stack, our solution achieves an average speedup of 1.8x (4x peak) at equal TOPS and memory resources across standard AI-benchmarks. Even against NPUs with double the compute and memory resources, Neutron delivers up to 3.3x higher performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。