光子芯片可大幅提升大模型算力与能效,突破电子硬件瓶颈。
What Is Next for LLMs? Next-Generation AI Computing Hardware Using Photonic Chips
- 用光子神经网络实现超高速矩阵运算,替代传统电子计算。
- 光子系统能效比电子芯片高数个数量级,适合超大规模模型训练。
- 适合研究下一代AI硬件的学者,尤其关注能效与长序列处理的团队。
大语言模型(LLMs)正快速逼近现有计算硬件的极限。例如,训练GPT-3估计耗电约1300兆瓦时,未来模型或需城市级(吉瓦级)电力预算。这促使人们探索超越传统冯·诺依曼架构的计算范式。本文综述面向下一代生成式AI的新型光子硬件,包括集成光子神经网络架构(如马赫-曾德尔干涉仪阵列、激光器、波长复用微环谐振器),可实现超快矩阵运算。还探讨了脉冲神经网络电路与混合自旋电子-光子突触等类脑器件,兼具存算功能。二维材料(石墨烯、过渡金属二硫属化物)在硅基光子平台中的应用被分析,用于可调调制器和片上突触元件。结合Transformer架构(自注意力与前馈层),分析其动态矩阵乘法映射到新型硬件的策略与挑战。进一步剖析ChatGPT、DeepSeek、LLaMA等主流模型的异同。综合当前先进组件、算法与集成技术,揭示规模化部署超大模型的关键进展与开放问题。结果表明,光子计算系统在吞吐量与能效方面有望较电子处理器提升数个数量级,但需突破长上下文窗口与超大数据集存储的内存瓶颈。
原文摘要 · Abstract (English)
Large language models (LLMs) are rapidly pushing the limits of contemporary computing hardware. For example, training GPT-3 has been estimated to consume around 1300 MWh of electricity, and projections suggest future models may require city-scale (gigawatt) power budgets. These demands motivate exploration of computing paradigms beyond conventional von Neumann architectures. This review surveys emerging photonic hardware optimized for next-generation generative AI computing. We discuss integrated photonic neural network architectures (e.g., Mach-Zehnder interferometer meshes, lasers, wavelength-multiplexed microring resonators) that perform ultrafast matrix operations. We also examine promising alternative neuromorphic devices, including spiking neural network circuits and hybrid spintronic-photonic synapses, which combine memory and processing. The integration of two-dimensional materials (graphene, TMDCs) into silicon photonic platforms is reviewed for tunable modulators and on-chip synaptic elements. Transformer-based LLM architectures (self-attention and feed-forward layers) are analyzed in this context, identifying strategies and challenges for mapping dynamic matrix multiplications onto these novel hardware substrates. We then dissect the mechanisms of mainstream LLMs, such as ChatGPT, DeepSeek, and LLaMA, highlighting their architectural similarities and differences. We synthesize state-of-the-art components, algorithms, and integration methods, highlighting key advances and open issues in scaling such systems to mega-sized LLM models. We find that photonic computing systems could potentially surpass electronic processors by orders of magnitude in throughput and energy efficiency, but require breakthroughs in memory, especially for long-context windows and long token sequences, and in storage of ultra-large datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。