用流式机器学习实时预测文件访问位置,提升存储系统预取效率。
Dynamic Adaptation in Data Storage: Real-Time Machine Learning for Enhanced Prefetching
- 采用流式分类模型动态预测文件访问偏移量。
- 在真实生产数据上实现更高预测准确率与内存效率。
- 适合需要实时响应的大型存储系统优化场景。
数据存储需求的指数级增长推动了分层存储管理策略的演进。本文探索将流式机器学习应用于多层级存储系统中的数据预取,以实现革新。与传统批处理训练模型不同,流式机器学习具备适应性强、实时洞察和计算高效的特点,能动态响应工作负载变化。本研究设计并验证了一个创新框架,集成流式分类模型以预测文件访问模式,特别是下一个文件偏移量。通过全面的特征工程和对大规模生产轨迹的实时评估,该方法在预测准确性、内存效率和系统适应性方面均取得显著提升。结果表明,流式模型在实时存储管理中具有巨大潜力,为先进的缓存与分层策略树立了新范式。
原文摘要 · Abstract (English)
The exponential growth of data storage demands has necessitated the evolution of hierarchical storage management strategies [1]. This study explores the application of streaming machine learning [3] to revolutionize data prefetching within multi-tiered storage systems. Unlike traditional batch-trained models, streaming machine learning [5] offers adaptability, real-time insights, and computational efficiency, responding dynamically to workload variations. This work designs and validates an innovative framework that integrates streaming classification models for predicting file access patterns, specifically the next file offset. Leveraging comprehensive feature engineering and real-time evaluation over extensive production traces, the proposed methodology achieves substantial improvements in prediction accuracy, memory efficiency, and system adaptability. The results underscore the potential of streaming models in real-time storage management, setting a precedent for advanced caching and tiering strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。