用光子硬件实现大规模神经网络训练,突破传统计算瓶颈。
Streamlined optical training of large-scale modern deep learning architectures with direct feedback alignment
- 采用直接反馈对齐算法,在光电混合平台完成大规模训练。
- 1500 TeraOPS算力下,成功训练超10亿参数的Transformer模型。
- 适合追求能效比与超大规模模型训练的研究者与工程团队。
现代深度学习几乎完全依赖专用电子硬件加速器。光子方法因低功耗和高速度,虽被广泛考虑用于推理,但至今仍主要局限于较简单任务。与此同时,基于反向传播训练深层复杂神经网络的问题,仍是当前架构规模与性能提升的主要限制,并造成重大计算与能耗瓶颈。本文在混合电子-光子平台上实验实现了通用且可扩展的训练算法——直接反馈对齐。光学处理单元执行大规模随机矩阵乘法,作为该算法的核心操作,在30瓦功耗下达到最高1500 TeraOPS的运算速度。我们实现了对包括超过10亿参数的Transformer在内的现代深度学习架构的光学训练,并在语言、视觉及基于扩散的生成任务中获得良好表现。研究了训练时间的缩放特性,证明该混合光电器件方法在超深宽神经网络上的潜在优势,为突破传统冯·诺依曼架构限制、持续推动现代人工智能指数级增长提供了新路径。
原文摘要 · Abstract (English)
Modern deep learning relies nearly exclusively on dedicated electronic hardware accelerators. Photonic approaches, with low consumption and high operation speed, are increasingly considered for inference but, to date, remain mostly limited to relatively basic tasks. Simultaneously, the problem of training deep and complex neural networks, overwhelmingly performed through backpropagation, remains a significant limitation to the size and, consequently, the performance of current architectures and a major compute and energy bottleneck. Here, we experimentally implement a versatile and scalable training algorithm, called direct feedback alignment, on a hybrid electronic-photonic platform. An optical processing unit performs large-scale random matrix multiplications, which is the central operation of this algorithm, at speeds up to 1500 TeraOPS under 30 Watts of power. We perform optical training of modern deep learning architectures, including Transformers, with more than 1B parameters, and obtain good performances on language, vision, and diffusion-based generative tasks. We study the scaling of the training time, and demonstrate a potential advantage of our hybrid opto-electronic approach for ultra-deep and wide neural networks, thus opening a promising route to sustain the exponential growth of modern artificial intelligence beyond traditional von Neumann approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。