arXiv:2511.11624cs.DCcs.AI2025-11被引 4

对比五种小模型在边缘设备的能耗与效率,找最优部署方案。

Characterizing and Understanding Energy Footprint and Efficiency of Small Language Model on Edges

  • 实测五种小模型在树莓派、Jetson等设备上的能效表现。
  • Jetson Orin Nano GPU版能效最高,Llama 3.2平衡性最佳。
  • 适合低功耗场景或高精度需求的用户参考选型。

基于云端的大语言模型及其变体已广泛应用于实际场景。将小型语言模型(SLMs)部署于边缘设备可降低延迟并摆脱网络依赖。然而,边缘设备受限的计算资源和能源预算对高效部署构成挑战。本研究评估了五种代表性SLMs——Llama 3.2、Phi-3 Mini、TinyLlama和Gemma 2——在Raspberry Pi 5、Jetson Nano及Jetson Orin Nano(CPU与GPU配置)上的能效表现。结果表明,搭载GPU加速的Jetson Orin Nano实现了最高的能量-性能比,显著优于基于CPU的设置。Llama 3.2在准确率与能效之间提供了最佳平衡,而TinyLlama虽精度较低但适合低功耗环境。相反,Phi-3 Mini尽管准确率高,却消耗最多能量。此外,GPU加速、内存带宽和模型架构是优化推理能效的关键因素。本研究通过实证分析,为人工智能、智能系统及移动自组网平台在能源受限环境下权衡准确率、推理延迟与能效提供实用指导。

原文摘要 · Abstract (English)

Cloud-based large language models (LLMs) and their variants have significantly influenced real-world applications. Deploying smaller models (i.e., small language models (SLMs)) on edge devices offers additional advantages, such as reduced latency and independence from network connectivity. However, edge devices' limited computing resources and constrained energy budgets challenge efficient deployment. This study evaluates the power efficiency of five representative SLMs - Llama 3.2, Phi-3 Mini, TinyLlama, and Gemma 2 on Raspberry Pi 5, Jetson Nano, and Jetson Orin Nano (CPU and GPU configurations). Results show that Jetson Orin Nano with GPU acceleration achieves the highest energy-to-performance ratio, significantly outperforming CPU-based setups. Llama 3.2 provides the best balance of accuracy and power efficiency, while TinyLlama is well-suited for low-power environments at the cost of reduced accuracy. In contrast, Phi-3 Mini consumes the most energy despite its high accuracy. In addition, GPU acceleration, memory bandwidth, and model architecture are key in optimizing inference energy efficiency. Our empirical analysis offers practical insights for AI, smart systems, and mobile ad-hoc platforms to leverage tradeoffs from accuracy, inference latency, and power efficiency in energy-constrained environments.

边缘计算能效分析小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。