对比主流边缘设备的推理性能与能耗,指导AI部署选型。
On the Sustainability of AI Inferences in the Edge
- 在Raspberry Pi、Jetson Nano等设备上测试多种模型推理表现。
- 发现模型精度、延迟、功耗与内存占用间存在可权衡的取舍关系。
- 适合物联网、智能工业等低延迟高能效场景的开发者参考。
物联网(IoT)及其前沿的AI应用(如自动驾驶和智能工业)融合了数据驱动系统与边缘部署。边缘设备通常执行推理以支持低延迟应用。除计算性能外,资源受限的边缘设备的能耗是实际部署的关键因素。典型设备包括Raspberry Pi(RPi)、Intel Neural Compute Stick(INCS)、NVIDIA Jetson nano(NJn)和Google Coral USB(GCU)。尽管这些设备被广泛用于边缘AI推理,但缺乏对其性能与能耗的系统性评估,难以支持设备与模型选型决策。本研究通过严格表征传统模型、神经网络及大语言模型在上述设备上的表现,分析模型F1分数、推理时间、推理功耗与内存使用间的权衡关系。硬件优化、框架调整及外部参数调优可协同平衡模型性能与资源消耗,助力实用化边缘AI部署。
原文摘要 · Abstract (English)
The proliferation of the Internet of Things (IoT) and its cutting-edge AI-enabled applications (e.g., autonomous vehicles and smart industries) combine two paradigms: data-driven systems and their deployment on the edge. Usually, edge devices perform inferences to support latency-critical applications. In addition to the performance of these resource-constrained edge devices, their energy usage is a critical factor in adopting and deploying edge applications. Examples of such devices include Raspberry Pi (RPi), Intel Neural Compute Stick (INCS), NVIDIA Jetson nano (NJn), and Google Coral USB (GCU). Despite their adoption in edge deployment for AI inferences, there is no study on their performance and energy usage for informed decision-making on the device and model selection to meet the demands of applications. This study fills the gap by rigorously characterizing the performance of traditional, neural networks, and large language models on the above-edge devices. Specifically, we analyze trade-offs among model F1 score, inference time, inference power, and memory usage. Hardware and framework optimization, along with external parameter tuning of AI models, can balance between model performance and resource usage to realize practical edge AI deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。