利用光电器件自然衰减实现高效脉冲变压器,显著降低能耗。
Otters++: A Time-to-first-spike Based Energy Efficient Optical Spiking Transformer

- 用In₂O₃器件衰减直接实现TTFS时间项,省去数字计算开销。
- 在GLUE上达84.17%平均得分,能耗优于已有脉冲变压器方案。
- 支持硬件真实噪声建模,适合低功耗部署的脉冲神经网络研究者。
脉冲神经网络(SNN)有望实现高能效推理,其中首次脉冲时间(TTFS)编码因每个神经元最多只发放一次而尤为吸引人。然而,实际中该优势常被计算时间衰减项和乘以突触权重的开销所抵消。为此,我们利用光电器件中的自然信号衰减这一物理‘缺陷’,将其作为TTFS的核心计算,命名为Otters++。具体地,采用定制的In₂O₃光电器件测量得到的衰减特性,直接实现TTFS的时间项,无需显式数字衰减计算。为将该思路扩展至Transformer模型,我们建立了Otters++与量化神经网络(QNN)之间的逐层功能等价关系,并开发了一种混合训练方法:前向传播使用设备真实的SNN计算,反向传播通过等效的QNN路径进行直通梯度传递,结合模型蒸馏。该方法避免了对离散首次脉冲事件的求导,缓解了直接TTFS-SNN训练中的过稀疏问题。此外,训练过程考虑了器件的运行间波动,系统级能耗模型也纳入了设备共享与多跳通信因素。在GLUE数据集上,Otters++平均得分提升至84.17%,同时保持明显能耗优势,表明基于物理实现的TTFS计算可在真实硬件条件下实现高效、可训练且鲁棒。
原文摘要 · Abstract (English)
Spiking neural networks (SNNs) are promising for energy-efficient inference, and time-to-first-spike (TTFS) coding is especially attractive because each neuron fires at most once. In practice, however, this benefit is often reduced by the cost of computing a temporal decay term and multiplying it by the synaptic weight. We address this issue by turning a physical hardware "bug," the natural signal decay in optoelectronic devices, into the main computation of TTFS, named Otters++. Specifically, we use the measured decay of a custom In$_2$O$_3$ optoelectronic synapse to directly realize the TTFS temporal term, removing the need for explicit digital decay computation. To scale this idea to Transformer models, we establish a layer-wise functional equivalence between the Otters++ and a quantized neural network (QNN), and develop a hybrid training method that uses device-faithful SNN computation in the forward pass and QNN straight-through gradients through the equivalent QNN path in the backward pass, together with model distillation. This avoids differentiation through discrete first-spike events and reduces the over-sparsity problem in direct TTFS-SNN training. We further make training aware of measured device noise by sampling run-to-run variation, and refine the system-level energy model by accounting for device sharing and multi-hop communication. On GLUE dataset, Otters++ improves the average score to 84.17\% while maintaining a clear energy advantage over prior spiking Transformer baselines. These results show that physically grounded TTFS computing can be efficient, trainable, and robust under realistic hardware effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。