探索小模型在边缘部署的性能与成本平衡,给出因地制宜的推理方案。
Edge-First Language Model Inference: Models, Metrics, and Tradeoffs
- 对比单设备与分布式边缘集群的推理能力,分析适用场景。
- 发现边缘推理在部分任务上可媲美云端,且成本更低。
- 适合关注低延迟、隐私保护的边缘AI系统设计者。
语言模型在各行业的广泛应用正推动其向计算连续体(从云到网络边缘)部署。这一转变旨在降低资源开销、减少延迟,并提升可靠性和隐私保护。得益于模型压缩技术的发展,小型语言模型(SLMs)成为边缘部署的核心,使资源受限的边缘设备实现本地推理成为可能。本文通过详尽的基准测试,评估了单个边缘设备上SLM的能力,并扩展至分布式边缘集群。研究发现,在某些场景下,边缘推理能实现与云端相当的性能且成本更低;而在模型容量或可扩展性受限时,云端回退则成为必要。本工作不提供统一解决方案,而是通过平台级对比和设计洞察,为构建高效、自适应的跨异构环境语言模型推理系统提供支持。
原文摘要 · Abstract (English)
The widespread adoption of Language Models (LMs) across industries is driving interest in deploying these services across the computing continuum, from the cloud to the network edge. This shift aims to reduce costs, lower latency, and improve reliability and privacy. Small Language Models (SLMs), enabled by advances in model compression, are central to this shift, offering a path to on-device inference on resource-constrained edge platforms. This work examines the interplay between edge and cloud deployments, starting from detailed benchmarking of SLM capabilities on single edge devices, and extending to distributed edge clusters. We identify scenarios where edge inference offers comparable performance with lower costs, and others where cloud fallback becomes essential due to limits in scalability or model capacity. Rather than proposing a one-size-fits-all solution, we present platform-level comparisons and design insights for building efficient, adaptive LM inference systems across heterogeneous environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。