arXiv:2501.10187cs.ARcs.AI2025-01被引 4

用小而廉价的Lite-GPU构建高效AI集群,提升可扩展性和能效。

Good things come in small packages: Should we build AI clusters with Lite-GPUs?

  • 用多个小型Lite-GPU组成集群,替代单一大型GPU。
  • 制造成本降低,故障影响范围更小,能效更高。
  • 适合大规模部署,尤其关注成本与稳定性的团队。

为应对生成式AI负载激增,传统做法是将更多算力和内存集成于单一复杂且昂贵的GPU中。然而,当前顶级GPU已面临封装、良率和散热瓶颈,其可扩展性存疑。本文提出重新思考AI集群设计:通过高效互联的大规模Lite-GPU集群实现扩展。Lite-GPU采用单个小芯片,性能仅为大GPU的一小部分。借助近期共封装光学技术,可实现多Lite-GPU间高带宽、低延迟通信,有效分担工作负载。本文分析了Lite-GPU在制造成本、故障传播范围、良率及能效方面的优势,并探讨了资源调度、任务分配、内存管理与网络架构等系统层面的机遇与挑战。

原文摘要 · Abstract (English)

To match the blooming demand of generative AI workloads, GPU designers have so far been trying to pack more and more compute and memory into single complex and expensive packages. However, there is growing uncertainty about the scalability of individual GPUs and thus AI clusters, as state-of-the-art GPUs are already displaying packaging, yield, and cooling limitations. We propose to rethink the design and scaling of AI clusters through efficiently-connected large clusters of Lite-GPUs, GPUs with single, small dies and a fraction of the capabilities of larger GPUs. We think recent advances in co-packaged optics can enable distributing AI workloads onto many Lite-GPUs through high bandwidth and efficient communication. In this paper, we present the key benefits of Lite-GPUs on manufacturing cost, blast radius, yield, and power efficiency; and discuss systems opportunities and challenges around resource, workload, memory, and network management.

AI集群Lite-GPU能效优化硬件设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。