arXiv:2605.29942physics.app-pheess.IV2026-05

用磁旋涡神经元+多状态存算一体芯片实现低功耗卷积网络

Reconfigurable Multistate MRAM Synapses with Vortex STNO based Neurons for Scalable In-Memory Convolutional Neural Networks

  • 将突触与神经元集成于单芯片,用磁场和电流调控多阻态
  • 在多个数据集上达到99.76%~56.46%准确率,每轮计算仅耗200皮焦
  • 适合超低功耗、可扩展的类脑芯片设计,尤其适合边缘部署

基于磁隧道结(MTJ)的磁随机存取存储器(MRAM)因其非易失性、高耐久性、快速切换速度和与CMOS工艺兼容性,是类脑计算和存内计算的有力候选。然而,传统的自旋转移矩和自旋轨道矩MRAM在神经网络应用中常面临高临界开关电流、大延迟、热不稳定性和显著的读写开销。本文展示了一种统一的多状态MRAM-自旋扭矩纳米振荡器(STNO)架构,将突触与神经元集成于单芯片,适用于卷积神经网络(CNN)应用。该系统采用1×8多状态MRAM阵列作为可编程突触,搭配基于涡旋结构的STNO神经元,通过场线驱动写入通道实现个体与集体编程。通过协同调节内部与外部磁场及偏置电流,实现多种可配置电阻状态,支持量化正负突触权重,以完成可配置的卷积核与池化操作。在MNIST、SVHN、CIFAR-10、Google语音命令(GSC)和RadioML数据集上的仿真验证显示,准确率分别达到99.76%、87.93%、78.14%、87.96%和56.46%。基于实际器件尺寸,整个架构占地约6171.2 μm²,MNIST任务下每训练/推理周期平均能耗为200.08 pJ,展现出其在可扩展低功耗类脑计算中的潜力。

原文摘要 · Abstract (English)

Magnetic tunnel junction (MTJ)-based magnetic random-access memory (MRAM) is a promising platform for neuromorphic and in-memory computing owing to its non-volatility, high endurance, fast switching dynamics and CMOS compatibility. However, conventional spin-transfer torque and spin-orbit torque MRAM implementations for neural networks often suffer from high critical switching currents, large latency, thermal instability and significant read-write overheads. Here, we demonstrate a unified multistate MRAM-spin-torque nano-oscillator (STNO) architecture that integrates synapses and neurons on a single chip for convolutional neural network (CNN) applications. The system employs 1x8 multistate MRAM arrays as programmable synapses coupled with a vortex-based STNO neuron, enabling both individual and collective programming through fieldline-driven write channels. Multiple configurable resistance states are achieved by tuning internal and external magnetic fields together with bias currents, allowing quantized positive and negative synaptic weights for configurable kernel and pooling operations. The proposed architecture is evaluated through simulation on MNIST, SVHN, CIFAR-10, Google Speech Commands (GSC) and RadioML datasets, achieving accuracy of 99.76%, 87.93%, 78.14%, 87.96% and 56.46% respectively. Based on fabricated device dimensions, the complete architecture occupies ~6171.2 μm2 with an average energy consumption of 200.08 pJ per training and inference cycle for MNIST, highlighting its potential for scalable low-power neuromorphic computing

存算一体类脑芯片磁存储低功耗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。