arXiv:2502.01670cs.ARcs.ET2025-02被引 11

用结构压缩提升光神经网络效率,降低硬件需求与功耗。

Hardware-Efficient Photonic Tensor Core: Accelerating Deep Neural Networks with Structured Compression

  • 采用块循环结构压缩光计算核心,减少模型与硬件开销。
  • 参数量降低74.91%,仍保持与原始模型相当的准确率。
  • 软硬件协同设计,提升光芯片非理想条件下的鲁棒性。

人工智能算力需求迅猛增长,传统电子硬件已难满足。光计算凭借并行性、高速度和低功耗优势成为潜在替代方案,但现有光子集成电路存在面积大、电光接口昂贵、控制复杂等问题,限制了光神经网络(ONNs)的实际可扩展性。为此,本文提出一种块循环光张量核心,构建结构压缩光神经网络(StrC-ONN)架构。该结构压缩技术显著降低模型复杂度与硬件资源消耗,同时保持神经网络的灵活性,且精度接近未压缩模型。此外,我们设计了一种硬件感知训练框架,以补偿片上非理想因素,提升模型鲁棒性与精度。实验在图像处理与分类任务中验证:StrC-ONN 可实现高达74.91%的可训练参数缩减,仍保持竞争力的准确率。性能分析表明,该软硬件协同设计方法预计可提升3.56倍能效。通过多维度降低硬件需求与控制复杂度,本工作为实用化、可扩展的光神经网络开辟新路径,有望应对未来计算效率挑战。

原文摘要 · Abstract (English)

The rapid growth in computing demands, particularly driven by artificial intelligence applications, has begun to exceed the capabilities of traditional electronic hardware. Optical computing offers a promising alternative due to its parallelism, high computational speed, and low power consumption. However, existing photonic integrated circuits are constrained by large footprints, costly electro-optical interfaces, and complex control mechanisms, limiting the practical scalability of optical neural networks (ONNs). To address these limitations, we introduce a block-circulant photonic tensor core for a structure-compressed optical neural network (StrC-ONN) architecture. The structured compression technique substantially reduces both model complexity and hardware resources without sacrificing the versatility of neural networks, and achieves accuracy comparable to uncompressed models. Additionally, we propose a hardware-aware training framework to compensate for on-chip nonidealities to improve model robustness and accuracy. Experimental validation through image processing and classification tasks demonstrates that our StrC-ONN achieves a reduction in trainable parameters of up to 74.91%,while still maintaining competitive accuracy levels. Performance analyses further indicate that this hardware-software co-design approach is expected to yield a 3.56 times improvement in power efficiency. By reducing both hardware requirements and control complexity across multiple dimensions, this work explores a new pathway toward practical and scalable ONNs, highlighting a promising route to address future computational efficiency challenges.

光计算神经网络能效优化硬件加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。