arXiv:2601.11935cs.DCcs.AI2026-01中稿 · and presented at t…被引 4

通过分析工作负载特征实现节能调度,降低云数据中心能耗15%-20%。

Big Data Workload Profiling for Energy-Aware Cloud Resource Management

  • 结合历史日志与实时数据,动态预测虚拟机部署的能效影响。
  • 在典型大数据任务上实现15%-20%的能耗降低,性能几乎无损。
  • 适合关注云资源节能、运维优化的研究者与工程师使用。

随着大数据工作负载规模与复杂性持续增长,云数据中心面临降低运营能耗的巨大压力。本文提出一种面向工作负载的节能调度框架,通过分析CPU利用率、内存需求和存储IO行为,指导虚拟机部署决策。该系统融合历史执行日志与实时遥测数据,预测候选部署方案的能效与性能影响,支持自适应资源合并,同时保障服务等级协议(SLA)合规性。在多节点云测试平台上,对典型Hadoop MapReduce、Spark MLlib和ETL工作负载进行评估。实验结果表明,相比基线调度器,该框架实现了15%至20%的稳定能耗节约,性能损耗可忽略不计。研究证明,工作负载特征分析是提升云上大数据处理环境可持续性的可行且可扩展策略。

原文摘要 · Abstract (English)

Cloud data centers face increasing pressure to reduce operational energy consumption as big data workloads continue to grow in scale and complexity. This paper presents a workload aware and energy efficient scheduling framework that profiles CPU utilization, memory demand, and storage IO behavior to guide virtual machine placement decisions. By combining historical execution logs with real time telemetry, the proposed system predicts the energy and performance impact of candidate placements and enables adaptive consolidation while preserving service level agreement compliance. The framework is evaluated using representative Hadoop MapReduce, Spark MLlib, and ETL workloads deployed on a multi node cloud testbed. Experimental results demonstrate consistent energy savings of 15 to 20 percent compared to a baseline scheduler, with negligible performance degradation. These findings highlight workload profiling as a practical and scalable strategy for improving the sustainability of cloud based big data processing environments.

云计算能耗优化大数据资源调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。