用可解释的机器学习优化固态硬盘的错误管理,实现更智能的存储架构设计。
Co-Design of Memory-Storage Systems for Workload Awareness with Interpretable Models
- 基于可解释机器学习建模,联合优化固态硬盘的错误管理与芯片工艺变异
- 在数千个数据中心固态硬盘上验证,支持跨代技术的持续数据驱动设计
- 能自动学习工作负载特征,适用于不同场景的存储系统优化
基于NAND或新兴存储器件的固态存储架构(SSD)在可靠性与性能方面面临根本性挑战。为同时实现这些目标,需对内存组件与固件级错误管理(EM)算法进行协同设计。本文提出一种面向系统的机器学习方法与建模框架,用于协同设计EM子系统,并融合缩放硅工艺带来的自然变异性。该模型通过统计可解释且直观可读的机器学习算法,分析了NAND内存组件与EM算法在闪存转换抽象层上的交互行为,覆盖合成压力测试(如stress-focused、JEDEC)与仿真工作负载(如YCSB等)。该通用化协同设计框架评估了跨越多个世代的数千个数据中心级SSD,实现了连续、整体、数据驱动的架构演进。此外,该框架还实现了对EM-工作负载域的表示学习,显著拓展了跨多种工作负载的架构设计空间。
原文摘要 · Abstract (English)
Solid-state storage architectures based on NAND or emerging memory devices (SSD), are fundamentally architected and optimized for both reliability and performance. Achieving these simultaneous goals requires co-design of memory components with firmware-architected Error Management (EM) algorithms for density- and performance-scaled memory technologies. We describe a Machine Learning (ML) for systems methodology and modeling for co-designing the EM subsystem together with the natural variance inherent to scaled silicon process of memory components underlying SSD technology. The modeling analyzes NAND memory components and EM algorithms interacting with comprehensive suite of synthetic (stress-focused and JEDEC) and emulation (YCSB and similar) workloads across Flash Translation abstraction layers, by leveraging a statistically interpretable and intuitively explainable ML algorithm. The generalizable co-design framework evaluates several thousand datacenter SSDs spanning multiple generations of memory and storage technology. Consequently, the modeling framework enables continuous, holistic, data-driven design towards generational architectural advancements. We additionally demonstrate that the framework enables Representation Learning of the EM-workload domain for enhancement of the architectural design-space across broad spectrum of workloads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。