arXiv:2411.14759cs.LGcs.AI2024-11被引 2

用在线学习提升数据冷热识别效率,降低资源开销。

Hammer: Towards Efficient Hot-Cold Data Identification via Online Learning

  • 基于在线学习动态追踪数据访问模式
  • 冷热分类准确率达90%,显著优于传统方法
  • 适合大规模存储系统优化场景

在大数据和云计算环境中,高效管理存储资源需要准确识别数据的“冷”与“热”状态。传统方法如基于规则的算法和早期人工智能技术难以应对动态工作负载,导致准确性低、适应性差且运维成本高。为此,我们提出一种基于在线学习策略的新方案,能够动态适应数据访问模式的变化,实现更高精度和更低运营成本。通过合成数据集和真实数据集的严格测试,该方法在冷热分类上达到90%的准确率,同时显著降低计算与存储开销。

原文摘要 · Abstract (English)

Efficient management of storage resources in big data and cloud computing environments requires accurate identification of data's "cold" and "hot" states. Traditional methods, such as rule-based algorithms and early AI techniques, often struggle with dynamic workloads, leading to low accuracy, poor adaptability, and high operational overhead. To address these issues, we propose a novel solution based on online learning strategies. Our approach dynamically adapts to changing data access patterns, achieving higher accuracy and lower operational costs. Rigorous testing with both synthetic and real-world datasets demonstrates a significant improvement, achieving a 90% accuracy rate in hot-cold classification. Additionally, the computational and storage overheads are considerably reduced.

存储优化在线学习冷热数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。