arXiv:2511.15950cs.DCcs.AI2025-11被引 2

超低延迟高能效大模型推理系统,支持多实例并发运行。

A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference

  • 垂直整合硬件+软件,288块加速卡协同工作
  • 115 peta-ops@4-bit,2.8ms/token延迟,功耗仅30kW
  • 适配3亿至700亿参数模型,适合企业级智能应用部署

一个垂直集成的端到端研究原型系统,结合288块NorthPole神经推理加速卡、离线训练算法、高性能运行时栈和容器化推理流水线,构建可扩展的云推理服务。该系统在18台2U服务器上实现115 peta-ops(4位整数精度)算力与3.7 PB/s内存带宽,功耗仅30 kW,总重730 kg,占地0.67 m²(42U机架)。系统可在2,048上下文长度下并行运行3个80亿参数的开源IBM Granite-3.3-8b-instruct模型,支持28名用户,单用户间令牌延迟为2.8 ms。系统具备可扩展、模块化和可重构特性,支持多种模型规模与上下文长度,适用于现有数据中心(云/本地)部署企业级智能体工作流。例如,可支持18个30亿参数模型实例或1个700亿参数模型实例。

原文摘要 · Abstract (English)

A vertically integrated, end-to-end, research prototype system combines 288 NorthPole neural inference accelerator cards, offline training algorithms, a high-performance runtime stack, and a containerized inference pipeline to deliver a scalable and efficient cloud inference service. The system delivers 115 peta-ops at 4-bit integer precision and 3.7 PB/s of memory bandwidth across 18 2U servers, while consuming only 30 kW of power and weighing 730 kg in a 0.67 m^2 42U rack footprint. The system can run 3 simultaneous instances of the 8-billion-parameter open-source IBM Granite-3.3-8b-instruct model at 2,048 context length with 28 simultaneous users and a per-user inter-token latency of 2.8 ms. The system is scalable, modular, and reconfigurable, supporting various model sizes and context lengths, and is ideal for deploying agentic workflows for enterprise AI applications in existing data center (cloud, on-prem) environments. For example, the system can support 18 instances of a 3-billion-parameter model or a single instance of a 70-billion-parameter model.

大模型推理低延迟能效优化系统集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。