arXiv:2508.16011cs.ETcs.AR2025-08被引 1

混合内存计算架构提升图神经网络训练效率

HePGA: A Heterogeneous Processing-in-Memory based GNN Training Accelerator

  • 根据图神经网络特点,将不同计算层映射到多种内存器件上
  • 能效比提升3.8倍,算力密度提升6.8倍,精度不损失
  • 适用于新兴Transformer模型推理加速,适合高性能计算场景

存内计算(PIM)架构为加速图神经网络(GNN)训练与推理提供了有前景的方案。然而,不同类型的PIM设备(如ReRAM、FeFET、PCM、MRAM和SRAM)在功耗、延迟、面积和非理想特性方面各有优劣。通过3D集成支持的异构多核架构可在单平台上融合多种PIM设备,实现节能高效的GNN训练。本文提出一种基于3D异构PIM的GNN训练加速器——HePGA。我们利用图神经网络各层及其计算核的特性,优化其在不同PIM设备及平面层级上的映射策略。实验表明,与现有PIM架构相比,HePGA在能效(TOPS/W)上最高提升3.8倍,在算力效率(TOPS/mm²)上最高提升6.8倍,且不牺牲GNN预测精度。最后,我们验证了HePGA在加速新兴Transformer模型推理方面的适用性。

原文摘要 · Abstract (English)

Processing-In-Memory (PIM) architectures offer a promising approach to accelerate Graph Neural Network (GNN) training and inference. However, various PIM devices such as ReRAM, FeFET, PCM, MRAM, and SRAM exist, with each device offering unique trade-offs in terms of power, latency, area, and non-idealities. A heterogeneous manycore architecture enabled by 3D integration can combine multiple PIM devices on a single platform, to enable energy-efficient and high-performance GNN training. In this work, we propose a 3D heterogeneous PIM-based accelerator for GNN training referred to as HePGA. We leverage the unique characteristics of GNN layers and associated computing kernels to optimize their mapping on to different PIM devices as well as planar tiers. Our experimental analysis shows that HePGA outperforms existing PIM-based architectures by up to 3.8x and 6.8x in energy-efficiency (TOPS/W) and compute efficiency (TOPS/mm2) respectively, without sacrificing the GNN prediction accuracy. Finally, we demonstrate the applicability of HePGA to accelerate inferencing of emerging transformer models.

图神经网络存内计算异构加速3D集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。