arXiv:2411.04814cs.LG2024-11

优化神经网络在非易失性存算交叉阵列上的映射布局

A Simple Packing Algorithm for Optimized Mapping of Artificial Neural Networks onto Non-Volatile Memory Cross-Bar Arrays

  • 提出简化算法自动确定最优物理阵列数量与尺寸
  • 发现最佳映射不依赖最少阵列数,而是阵列容量与外围电路协同作用
  • 揭示方形阵列未必最优,性能提升常需更大总面积

存算交叉阵列已成为提升机器学习计算效率的有前景方案。以往研究主要聚焦于实现基本数学运算。本文探索将人工神经网络各层映射到芯片上分块排列的物理交叉阵列中的影响。我们提出一种简化映射算法,用于确定固定最优阵列尺寸下的物理阵列数量,并估算满足给定设计目标时所需最小面积。该算法与传统的二元线性优化方法(解决等效装箱问题)对比。结果表明,最优解并不必然对应最少阵列数,而是由阵列容量与外围电路缩放特性之间的相互作用决定。此外,我们发现方形阵列并非始终为最优选择,性能优化往往以增加总阵列面积为代价。

原文摘要 · Abstract (English)

Neuromorphic computing with crossbar arrays has emerged as a promising alternative to improve computing efficiency for machine learning. Previous work has focused on implementing crossbar arrays to perform basic mathematical operations. However, in this paper, we explore the impact of mapping the layers of an artificial neural network onto physical cross-bar arrays arranged in tiles across a chip. We have developed a simplified mapping algorithm to determine the number of physical tiles, with fixed optimal array dimensions, and to estimate the minimum area occupied by these tiles for a given design objective. This simplified algorithm is compared with conventional binary linear optimization, which solves the equivalent bin-packing problem. We have found that the optimum solution is not necessarily related to the minimum number of tiles; rather, it is shown to be an interaction between tile array capacity and the scaling properties of its peripheral circuits. Additionally, we have discovered that square arrays are not always the best choice for optimal mapping, and that performance optimization comes at the cost of total tile area

存算一体神经网络布局优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。