用磁隧道结内存实现并行随机计算,减少数据移动与能耗
Maximizing Memory-Level Parallelism via Integrated Stochastic Logic-in-Memory Architectures

- 将随机计算嵌入磁性内存,直接在存储单元内生成比特流
- 无需外部随机数生成电路,支持并行计算核心函数且抗噪声
- 适合低功耗高并发的边缘计算场景
当前高性能架构受数据移动延迟和能量开销制约,单核性能提升放缓与高度数据密集型任务并行出现。存内计算架构通过缓解内存带宽瓶颈、利用大规模并发、减少存储与计算单元间数据移动,成为传统冯·诺依曼系统的补充方案。本文提出一种并行存内随机计算(SC)架构,在具备逻辑内建(LIM)能力的磁隧道结(MTJ)内存中实现端到端计算流水线。利用MTJ器件固有的随机性与写读特性,该架构可完全并行且确定性地将二进制操作数转换为概率比特流,避免了高能耗的外部随机数生成电路。这些比特流由集成在内存阵列内的并行随机算术单元处理,以极低硬件复杂度高效实现核心算术与超越函数运算,并具有固有噪声容忍性。所得随机输出可作为后续随机处理输入,或通过并行累加机制转换回二进制形式并存回MTJ内存。通过在统一存内结构中紧密集成数据存储、比特流生成与计算,该设计最大化内存级并行性,显著降低数据移动。
原文摘要 · Abstract (English)
Today's high-performance architectures are increasingly constrained by data movement latency and energy overhead, as the slowdown of single-core performance scaling coincides with the rise of highly data-intensive workloads. In-memory architectures have emerged as a complementary solution to conventional von Neumann systems by alleviating memory bandwidth bottlenecks, exploiting massive concurrency, and mitigating excessive data movement between memory and processing units. This study proposes a parallel in-memory stochastic computing (SC) architecture that implements an end-to-end computation pipeline within Magnetic Tunnel Junction (MTJ)-based memory augmented with logic-in-memory (LIM) capabilities. By leveraging the inherent stochasticity and write-read characteristics of MTJ devices, the proposed architecture enables a fully parallel and deterministic conversion of binary operands into probabilistic bit-streams, eliminating the need for energy-intensive external random number generation circuitry. These bit-streams are processed by parallel stochastic arithmetic units integrated directly within the memory arrays to efficiently implement core arithmetic and transcendental functions with minimal hardware complexity and inherent noise tolerance. The resulting stochastic outputs can be either reused as an input of future stochastic processing or converted back to binary form using parallel accumulation mechanisms and stored in the MTJ memory. By tightly integrating data storage, bit-stream generation, and computation within a unified in-memory fabric, the proposed design maximizes memory-level parallelism while substantially minimizing data movement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。