评测边缘设备上大模型推理性能,助力隐私敏感场景落地
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers

- 构建多维评测框架,对比四类边缘硬件加速器表现
- 实测显示专用加速器提升吞吐量并降低功耗
- 适合无人机、野外设备等低带宽高安全需求场景
大型语言模型(LLMs)在小参数规模下能力持续提升。传统云部署面临数据隐私、延迟和成本问题,尤其在工控与国防场景中尤为突出。模型压缩、量化及低成本边缘加速器的发展使单板计算机本地推理成为可能,但配置空间维度高,缺乏系统评估难以找到最优方案。现有边缘基准测试多依赖纯CPU、覆盖设备有限,且评估任务单一,缺乏对硬件效能的多维分析。本文提出一种多维度基准方法,联合评估四种适配物联网的边缘平台在推理性能与硬件效率上的表现,涵盖最新硬件加速器的单板计算机。结果表明,NPU与GPU等专用加速器显著提升能效与每秒生成词元数(token throughput),同时量化了功耗、设备尺寸与吞吐量间的权衡关系。研究为无人设备、便携式坚固化操作等隐私敏感、连接受限环境中的生成式AI部署提供实用指导。
原文摘要 · Abstract (English)
Large language models (LLMs) are becoming increasingly capable at small parameter scales. At the same time, conventional cloud-centric deployment introduces challenges around data privacy, latency, and cost that are acute in operational technology and defence environments. Advances in model distillation, quantisation, and affordable edge accelerators now make local LLM inference on single-board computers feasible, but the high dimensionality of the configuration space makes identifying optimal deployments difficult without structured evaluation. Existing LLM-specific edge benchmarking efforts rely on CPU-only inference, poor coverage of genuine single-board computers, and generic evaluation tasks that lack multi-dimensional assessment of hardware effectiveness. This paper proposes a multi-dimensional benchmarking methodology that jointly evaluates inference performance and hardware efficiency across four IoT-suitable edge platform configurations testing single-board computers with the latest available hardware accelerators. Our results reveal the benefits of using hardware accelerators such as NPUs and GPUs, along with multi-dimensional evaluations quantifying the trade-offs between power efficiency, physical device size and token throughput; offering practical guidance for deploying generative AI in privacy-sensitive and connectivity-limited environments such as unmanned vehicles and portable, ruggedised operations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。