arXiv:2507.14000cs.PFcs.AI2025-07被引 4

用光子芯片突破内存瓶颈,让AI训练推理更快更省电。

Photonic Fabric Platform for AI Accelerators

  • 光子芯片+高带宽内存构建2.5D系统,实现32TB共享内存与115Tbps交换
  • 大模型推理速度提升7倍,训练能耗降低60%-90%(1T参数)
  • 适用于所有受限于内存-算力固定比的AI加速器设计

本文提出Photonic FabricTM与Photonic Fabric ApplianceTM(PFA),一种基于光子技术的交换与内存子系统,具备低延迟、高带宽和每比特低能耗特性。通过在2.5D电光封装中集成高性能HBM3E内存、板载光子交换机及外部DDR5,PFA提供高达32TB的共享内存和115Tbps的全互连数字交换能力。该架构使分布式AI训练与推理能更高效执行并行策略,突破了现有几乎所有XPU加速器中固定的内存-计算比限制。将XPU上的本地HBM堆叠替换为连接光子织网的芯片小片,可大幅提升内存容量与带宽,实现超越单片封装HBM的弹性扩展。文中引入CelestiSim轻量级分析模拟器,在NVIDIA H100与H200系统上验证其性能。仿真显示,405B参数大模型推理吞吐提升3.66倍、延迟改善1.40倍;1T参数模型可达7.04倍吞吐提升与1.41倍延迟优化;所有大模型训练场景中数据搬运能耗降低60%-90%。尽管结果基于NVIDIA GPU,但可推广至其他存在相同内存-算力比例限制的AI加速器(XPUs)。

原文摘要 · Abstract (English)

This paper presents the Photonic FabricTM and the Photonic Fabric ApplianceTM (PFA), a photonic-enabled switch and memory subsystem that delivers low latency, high bandwidth, and low per-bit energy. By integrating high-bandwidth HBM3E memory, an on-module photonic switch, and external DDR5 in a 2.5D electro-optical system-in-package, the PFA offers up to 32 TB of shared memory alongside 115 Tbps of all-to-all digital switching. The Photonic FabricTM enables distributed AI training and inference to execute parallelism strategies more efficiently. The Photonic Fabric removes the silicon beachfront constraint that limits the fixed memory-to-compute ratio observed in virtually all current XPU accelerator designs. Replacing a local HBM stack on an XPU with a chiplet that connects to the Photonic Fabric increases its memory capacity and correspondingly its memory bandwidth by offering a flexible path to scaling well beyond the limitations of on-package HBM alone. We introduce CelestiSim, a lightweight analytical simulator validated on NVIDIA H100 and H200 systems. It is used to evaluate the performance of LLM reference and energy savings on PFA, without any significant change to the GPU core design. With the PFA, the simulation results show that up to 3.66x throughput and 1.40x latency improvements in LLM inference at 405B parameters, up to 7.04x throughput and 1.41x latency improvements at 1T parameters, and 60-90% energy savings in data movement for heavy collective operations in all LLM training scenarios. While these results are shown for NVIDIA GPUs, they can be applied similarly to other AI accelerator designs (XPUs) that share the same fundamental limitation of fixed memory to compute.

光子计算AI加速内存瓶颈大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。