arXiv:2507.02871cs.DCcs.AI2025-07

ZettaLith架构让AI推理效率提升千倍,功耗降低千倍以上。

ZettaLith: An Architectural Exploration of Extreme-Scale AI Inference Acceleration

  • 专为推理优化,通过多项协同设计实现极致能效。
  • 2027年单机柜可达1.507泽弗拉普斯,比当前GPU系统快1047倍。
  • 适合追求极致算力与能效的推理部署场景。

当前及未来AI系统的高计算成本和高功耗成为广泛部署与进一步扩展的主要障碍。现有硬件方法面临根本性效率瓶颈。本文提出ZettaLith,一种可扩展计算架构,相比当前基于GPU的系统,可将AI推理的成本和功耗降低超过1,000倍。基于架构分析与技术预测,一个单个ZettaLith机柜在2027年或可实现1.507泽弗拉普斯(zettaFLOPS)的性能——相较于当前领先GPU机柜,在FP4 Transformer推理上,理论性能提升达1,047倍,功耗效率提高1,490倍,成本效益提升2,325倍。该架构通过放弃通用GPU应用,并利用大量协同设计的架构创新,以成熟的数字电子技术实现上述突破。ZettaLith核心原则可高效缩放至艾弗拉普斯级桌面系统与拍弗拉普斯级移动芯片,均保持约1,000倍优势。其系统架构比当前复杂GPU集群更简洁,专为AI推理设计,不适用于训练。

原文摘要 · Abstract (English)

The high computational cost and power consumption of current and anticipated AI systems present a major challenge for widespread deployment and further scaling. Current hardware approaches face fundamental efficiency limits. This paper introduces ZettaLith, a scalable computing architecture designed to reduce the cost and power of AI inference by over 1,000x compared to current GPU-based systems. Based on architectural analysis and technology projections, a single ZettaLith rack could potentially achieve 1.507 zettaFLOPS in 2027 - representing a theoretical 1,047x improvement in inference performance, 1,490x better power efficiency, and could be 2,325x more cost-effective than current leading GPU racks for FP4 transformer inference. The ZettaLith architecture achieves these gains by abandoning general purpose GPU applications, and via the multiplicative effect of numerous co-designed architectural innovations using established digital electronic technologies, as detailed in this paper. ZettaLith's core architectural principles scale down efficiently to exaFLOPS desktop systems and petaFLOPS mobile chips, maintaining their roughly 1,000x advantage. ZettaLith presents a simpler system architecture compared to the complex hierarchy of current GPU clusters. ZettaLith is optimized exclusively for AI inference and is not applicable for AI training.

AI推理架构创新能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。