arXiv:2607.26016cs.ARcs.AI2026-07

用光子模式并行加速Transformer,功耗降低63.6%。

MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar

论文配图:MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar
图 1 · 摘自论文原文
  • 用模式复用光干涉实现矩阵运算,每根波导四路并行。
  • 实测功耗降63.6%,面积减40.4%,能效提升40.6%。
  • 兼容单激光连续波,适合部署在高能效场景。

近年来,光子Transformer加速器(PTA)已在加速Transformer推理方面显著优于电子加速器,展现出速度与能效优势。然而,现有方案依赖昂贵的多波长光源和带主动相位调制器的大型点积单元,导致效率低、不实用。为此,我们提出MDTransformer,一种基于模式复用光数据流与逆向设计的软硬件协同光子加速器。其核心为紧凑的模式复用光子张量核(MPTC),通过空间模式干涉实现复数矩阵运算,利用逆向设计的多模耦合器、交叉结构及马赫-曾德尔IQ调制器,每个导波模式(如TE0-TE3)作为独立计算通道,实现无需频谱滤波的四倍波导并行。结合相干探测与IQ调制,可编码幅度与相位,支持全范围复数运算。系统具备亚4比特有效精度,模间串扰低于-30 dB。其逆向设计支持1550 nm单激光连续波运行,具备可扩展性。实验表明,相比现有最优PTA,MDTransformer在DeiT-Tiny/Small/Base与BERT-Base/Large等负载下,实现40.4%面积缩减、63.6%功耗降低、40.6%能效提升,且延迟相当。结果表明,该方案为高性能、低功耗的Transformer系统提供了实用路径。

原文摘要 · Abstract (English)

Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficiency improvements over electronic accelerators for expediting Transformer inference. However, state-of-the-art rely on expensive multi-wavelength light generation and large dot-product units due to active phase-shifter components, thus making their approach inefficient and impractical. To address this, we propose MDTransformer, a novel hardware-software co-design of PTA based on mode-division optical dataflow and operations. Specifically, MDTransformer performs complex matrix operations using spatial-mode interference, that leverages the inverse-designed multi-mode couplers, crossings, and Mach-Zehnder IQ modulators into a compact mode-division photonic tensor core (MPTC), capable of executing matrix multiplications in the optical domain. Its each guided mode (i.e., TE0-TE3) acts as an independent computational lane, enabling four-fold parallelism-per-waveguide without spectral filtering or free-spectral-range limitations. Moreover, its coherent detection and IQ modulation jointly encode amplitude and phase, realizing complex-valued arithmetic for full-range operations in transformers. MDTransformer offers analog multiplication with sub-4-bit effective precision and inter-modal crosstalk below -30 dB. Its inverse-designed approach also offers scalable and full compatibility with single-laser continuous-wave operation at 1550 nm. Experimental results show that MDTransformer achieves 40.4% area reduction, 63.6% power saving, 40.6% energy saving, and comparable latency over the state-of-the-art PTA across different workloads (i.e., DeiT-Tiny/Small/Base and BERT-Base/Large). These results show that MDTransformer offers a practical solution for high-performance and energy-efficient transformer-based systems.

光子计算Transformer加速能效优化模式复用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。