在超算中心部署生成式AI服务,打通云与超算的容器化能力。
Experience Deploying Containerized GenAI Services at an HPC Center
- 构建融合HPC与Kubernetes的统一容器架构,支持多运行时部署。
- 成功在两种平台用vLLM部署Llama模型,实现跨环境可复现。
- 为超算社区提供实操指南,推动GenAI工具链发展。
生成式人工智能(GenAI)应用由推理服务器、对象存储、向量与图数据库及用户界面等专用组件构成,通过基于Web的API相互连接。尽管这些组件通常被容器化并在云环境中部署,但此类能力在高性能计算(HPC)中心仍处于初步阶段。本文分享了在成熟HPC中心部署GenAI工作负载的实际经验,探讨了HPC与云计算环境的集成。我们描述了一种融合HPC与Kubernetes平台的汇聚式计算架构,支持容器化GenAI工作负载的运行,提升可复现性。案例研究展示了使用容器化推理服务器vLLM,在Kubernetes和HPC平台均部署Llama大语言模型(LLM),并采用多种容器运行时。我们的经验突出了实际考量与机遇,为HPC容器社区提供了未来研究与工具开发的指导。
原文摘要 · Abstract (English)
Generative Artificial Intelligence (GenAI) applications are built from specialized components -- inference servers, object storage, vector and graph databases, and user interfaces -- interconnected via web-based APIs. While these components are often containerized and deployed in cloud environments, such capabilities are still emerging at High-Performance Computing (HPC) centers. In this paper, we share our experience deploying GenAI workloads within an established HPC center, discussing the integration of HPC and cloud computing environments. We describe our converged computing architecture that integrates HPC and Kubernetes platforms running containerized GenAI workloads, helping with reproducibility. A case study illustrates the deployment of the Llama Large Language Model (LLM) using a containerized inference server (vLLM) across both Kubernetes and HPC platforms using multiple container runtimes. Our experience highlights practical considerations and opportunities for the HPC container community, guiding future research and tool development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。