为AI推理设计新型内存,提升密度与读取带宽。
Managed-Retention Memory: A New Class of Memory for the AI Era
- 通过放弃长期保存和写入性能,优化读取带宽和密度
- 相比HBM,MRM在关键指标上表现更优,适合AI推理场景
- 适合关注低功耗、高密度内存的AI系统研发人员
当前AI集群是高带宽内存(HBM)的主要应用之一。然而,HBM在多个方面不适用于AI工作负载:其写入性能过度配置,而密度和读取带宽不足,且每比特能耗较高。此外,由于制造复杂,其良率低于DRAM,成本也更高。本文提出一种新型内存——管理保留内存(Managed-Retention Memory, MRM),专为存储AI推理中的关键数据结构而优化。我们认为,MRM或能为最初用于支持存储类内存(SCM)的技术提供可行性路径。这些技术传统上具备10年以上持久性,但存在输入输出性能差和耐久性不足的问题。MRM做出不同权衡,通过分析工作负载的IO模式,主动放弃长期数据保留和写入性能,以换取对AI工作负载更重要的性能优势。
原文摘要 · Abstract (English)
AI clusters today are one of the major uses of High Bandwidth Memory (HBM). However, HBM is suboptimal for AI workloads for several reasons. Analysis shows HBM is overprovisioned on write performance, but underprovisioned on density and read bandwidth, and also has significant energy per bit overheads. It is also expensive, with lower yield than DRAM due to manufacturing complexity. We propose a new memory class: Managed-Retention Memory (MRM), which is more optimized to store key data structures for AI inference workloads. We believe that MRM may finally provide a path to viability for technologies that were originally proposed to support Storage Class Memory (SCM). These technologies traditionally offered long-term persistence (10+ years) but provided poor IO performance and/or endurance. MRM makes different trade-offs, and by understanding the workload IO patterns, MRM foregoes long-term data retention and write performance for better potential performance on the metrics important for these workloads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。