arXiv:2504.20101cs.DCcs.AI2025-04被引 3

用去中心化网络降低大模型部署门槛,提升服务效率与隐私保护。

PlanetServe: A Decentralized, Scalable, and Privacy-Preserving Overlay for Democratizing Large Language Model Serving

  • 构建基于节点共享的去中心化服务网络,实现资源协同调度。
  • 相比基线方案,延迟降低超过50%,且安全机制开销极小。
  • 适合个人和小团队低成本部署与测试大模型应用。

尽管开源和低成本大语言模型(LLMs)在研究与开发方面已取得显著进展,但服务可扩展性仍是关键挑战,尤其对希望部署和测试其创新的小组织和个人而言。受点对点网络启发,我们提出GenTorrent,一种利用分散贡献者计算资源的大模型服务覆盖网络。本文系统分析了四个核心问题:1)覆盖网络组织;2)大模型通信隐私;3)覆盖转发以提高资源效率;4)服务品质验证。这是首个针对去中心化大模型服务中这些基础问题的系统性研究。原型在分布式节点上的评估表明,GenTorrent相比无覆盖转发的基线设计,延迟降低超过50%。同时,安全特性对服务延迟和吞吐量的影响极小。我们认为本工作为未来人工智能服务的民主化与规模化开辟了新方向。

原文摘要 · Abstract (English)

While significant progress has been made in research and development on open-source and cost-efficient large-language models (LLMs), serving scalability remains a critical challenge, particularly for small organizations and individuals seeking to deploy and test their LLM innovations. Inspired by peer-to-peer networks that leverage decentralized overlay nodes to increase throughput and availability, we propose GenTorrent, an LLM serving overlay that harnesses computing resources from decentralized contributors. We identify four key research problems inherent to enabling such a decentralized infrastructure: 1) overlay network organization; 2) LLM communication privacy; 3) overlay forwarding for resource efficiency; and 4) verification of serving quality. This work presents the first systematic study of these fundamental problems in the context of decentralized LLM serving. Evaluation results from a prototype implemented on a set of decentralized nodes demonstrate that GenTorrent achieves a latency reduction of over 50% compared to the baseline design without overlay forwarding. Furthermore, the security features introduce minimal overhead to serving latency and throughput. We believe this work pioneers a new direction for democratizing and scaling future AI serving capabilities.

大模型服务去中心化隐私保护资源调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。