Maia 200用软件定义数据流架构,实现超高效能大模型推理。
Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration

- 通过软件定义数据流,精准控制内存与数据移动引擎。
- 750W功耗下达14.5 Pflop/s FP4与5.072 Pflop/s FP8算力。
- 适合追求极致能效与大规模并行的AI推理系统设计者。
我们提出Maia 200,一种先进AI加速器,在750W热设计功耗和7TB/s HBM带宽下,实现14.5 Pflop/s FP4与5.072 Pflop/s FP8算力。Maia代表新型软件定义本地访问数据流架构(SDLA),通过显式编程数据流引擎,协调专用内存与数据搬运单元。该架构从传统线程中心转向数据移动中心,显著提升效率与可扩展性。受Flynn分类启发,我们构建了数据管理新范式,有效应对现代AI计算挑战。Maia 200在支持大规模并行推理的同时,大幅降低能耗与成本,是下一代高性能计算系统的有力候选方案。
原文摘要 · Abstract (English)
We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and 7 TB/s HBM bandwidth. Maia exemplifies a new class of Software Defined Locally Accessed Dataflow Architectures (SDLA), which explicitly program dataflow engines to orchestrate highly specialized memories and data movement engines. This approach shifts the focus from today's thread-centric to data-movement-centric architecture, improving efficiency and scalability. Our taxonomy of data management, inspired by Flynn's classification, highlights how SDLA addresses challenges in modern AI computing. Maia 200 achieves significant cost and energy savings while supporting massive parallelism for AI inference workloads, making it a compelling solution for next-generation high-performance computing systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。