用闲置显卡搭建去中心化大模型推理系统,降低成本。
DeServe: Towards Affordable Offline LLM Inference via Decentralization
- 利用空闲显卡资源,构建去中心化推理网络。
- 在高延迟网络下,吞吐量提升6.7至12.6倍。
- 适合预算有限、需离线部署的开发者使用。
生成式AI的快速发展及其在日常工作中日益广泛的应用,显著增加了对大语言模型(LLM)推理服务的需求。尽管专有模型仍占主流,但开源LLM的最新进展使其成为有力竞争者。然而,部署这些模型常受制于高成本和有限的GPU资源。为此,本文提出一种去中心化离线推理系统的架构设计。通过利用闲置的GPU资源,所提出的DeServe系统以更低的成本实现对LLM的分布式访问。DeServe特别针对高延迟网络环境下的推理吞吐量优化问题。实验表明,在此类条件下,DeServe相较现有基准系统实现了6.7倍至12.6倍的吞吐量提升。
原文摘要 · Abstract (English)
The rapid growth of generative AI and its integration into everyday workflows have significantly increased the demand for large language model (LLM) inference services. While proprietary models remain popular, recent advancements in open-source LLMs have positioned them as strong contenders. However, deploying these models is often constrained by the high costs and limited availability of GPU resources. In response, this paper presents the design of a decentralized offline serving system for LLM inference. Utilizing idle GPU resources, our proposed system, DeServe, decentralizes access to LLMs at a lower cost. DeServe specifically addresses key challenges in optimizing serving throughput in high-latency network environments. Experiments demonstrate that DeServe achieves a 6.7x-12.6x improvement in throughput over existing serving system baselines in such conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。