在树莓派上实现生成式AI边缘推理,验证轻量模型本地部署可行性。
Generative AI on the Edge: Architecture and Performance Evaluation
- 用树莓派5集群跑多种大语言模型,结合K3s轻量编排系统。
- 轻量模型(如Yi、Phi)在纯CPU下每秒生成5-12个词,资源占用低于50%。
- 适合偏远或带宽受限地区,无需依赖云端的本地智能服务。
6G网络原生愿景要求将先进智能嵌入网络并贴近用户,亟需对生成式AI(GenAI)在边缘设备上的表现进行系统评估。基于开放无线接入网(ORAN)和‘网络即盒子’的新兴方案,普遍倡导使用低成本、现成组件以简化部署,例如在农村连通性场景中。然而,针对通用边缘设备上大语言模型(LLM)的架构设计、硬件测试平台及精确性能量化仍缺乏研究。本研究在单个商品级树莓派上构建边缘测试床,探索各类大小模型在树莓派5集群上的计算密集型推理能力,采用轻量级Kubernetes发行版(K3s)与模块化提示机制。通过分析吞吐量、延迟、准确率与效率,验证其可行性与局限性。结果表明,仅用CPU部署轻量模型(如Yi、Phi、Llama3)即可有效支撑边缘应用,生成吞吐率达每秒5至12个词,且CPU与内存占用均低于50%。结论指出,生成式AI在边缘可实现远程或带宽受限环境下的本地化推理,无需依赖云基础设施。
原文摘要 · Abstract (English)
6G's AI native vision of embedding advance intelligence in the network while bringing it closer to the user requires a systematic evaluation of Generative AI (GenAI) models on edge devices. Rapidly emerging solutions based on Open RAN (ORAN) and Network-in-a-Box strongly advocate the use of low-cost, off-the-shelf components for simpler and efficient deployment, e.g., in provisioning rural connectivity. In this context, conceptual architecture, hardware testbeds and precise performance quantification of Large Language Models (LLMs) on off-the-shelf edge devices remains largely unexplored. This research investigates computationally demanding LLM inference on a single commodity Raspberry Pi serving as an edge testbed for ORAN. We investigate various LLMs, including small, medium and large models, on a Raspberry Pi 5 Cluster using a lightweight Kubernetes distribution (K3s) with modular prompting implementation. We study its feasibility and limitations by analyzing throughput, latency, accuracy and efficiency. Our findings indicate that CPU-only deployment of lightweight models, such as Yi, Phi, and Llama3, can effectively support edge applications, achieving a generation throughput of 5 to 12 tokens per second with less than 50\% CPU and RAM usage. We conclude that GenAI on the edge offers localized inference in remote or bandwidth-constrained environments in 6G networks without reliance on cloud infrastructure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。