将vLLM与K8s、Slurm融合,实现超算上大模型动态推理的高效扩展。
Automated Dynamic AI Inference Scaling on HPC-Infrastructure: Integrating Kubernetes, Slurm and vLLM
- 通过整合vLLM、K8s和Slurm,在超算上构建动态推理框架。
- 支持100至1000并发请求,端到端延迟仅增加约500毫秒。
- 适合需要弹性部署大模型的高校科研与高并发服务场景。
随着人工智能推理需求上升,尤其是在高等教育领域,利用现有基础设施的新方案不断涌现。高性能计算(HPC)已成为实现此类解决方案的主流方式。然而,传统HPC运行模式难以适应同步、面向用户的动态AI应用负载需求。本文提出一种基于超算系统RAMSES的解决方案,通过集成vLLM、Slurm与Kubernetes服务大模型推理。初步基准测试显示,该架构在100、500和1000并发请求下均能高效扩展,端到端延迟仅增加约500毫秒。
原文摘要 · Abstract (English)
Due to rising demands for Artificial Inteligence (AI) inference, especially in higher education, novel solutions utilising existing infrastructure are emerging. The utilisation of High-Performance Computing (HPC) has become a prevalent approach for the implementation of such solutions. However, the classical operating model of HPC does not adapt well to the requirements of synchronous, user-facing dynamic AI application workloads. In this paper, we propose our solution that serves LLMs by integrating vLLM, Slurm and Kubernetes on the supercomputer \textit{RAMSES}. The initial benchmark indicates that the proposed architecture scales efficiently for 100, 500 and 1000 concurrent requests, incurring only an overhead of approximately 500 ms in terms of end-to-end latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。