评测大模型在云资源调度中的实际效果,发现精准预测不等于好决策。
CloudCons: A Comprehensive End-to-End Benchmark for Cloud Resource Consolidation

- 构建端到端云资源合并评估基准,覆盖多家云平台真实负载。
- 发现大模型预测准但未必能提升资源调度效率,需调优预测分位数。
- 给出量化建议:如何平衡资源利用率与服务可靠性,适合运维工程师参考。
由于为保障服务可靠性而过度配置资源,云计算数据中心的资源利用率长期偏低。为缓解此问题,基于预测-优化范式的方法通过预判未来需求来优化资源合并。尽管新兴的时间序列基础模型有望通过零样本泛化能力提升该范式,但现有基准仅关注预测误差,未验证其对下游决策的实际效用,导致其真实价值存疑。为此,我们提出 CloudCons,一个面向云资源合并的综合性端到端评估基准。构建了涵盖华为云、微软 Azure 与 Google Borg 多样工作负载的高质量数据集,捕捉从同步昼夜节律到随机脉冲式突发及高频噪声等不同服务特征。对统计方法、深度学习模型与基础模型进行广泛评估。实验揭示关键发现:虽然基础模型在零样本预测中表现更优,但该优势并未自动转化为更好的决策效用。具有实际意义的是,我们系统分析了预测分位数选择作为关键调控杠杆的作用。提供可操作指南,用于校准分位数以平衡资源效率与服务可靠性,为实际部署提供重要依据。
原文摘要 · Abstract (English)
Driven by conservative over-provisioning to guarantee service reliability, resource utilization in cloud data centers remains at low levels. To mitigate this, the forecast-then-optimize paradigm has emerged to optimize consolidation by anticipating future demands. While emerging time series foundation models promise to enhance this paradigm through zero-shot generalization, existing benchmarks focus solely on prediction error metrics. The actual decision utility of these advanced models remains unverified, rendering their practical value for downstream tasks uncertain. To bridge this gap, we propose CloudCons, a comprehensive end-to-end benchmark designed to evaluate forecasting models within the specific context of cloud resource consolidation. We build high-quality datasets that cover diverse workloads from Huawei Cloud, Microsoft Azure, and Google Borg, capturing distinct service characteristics ranging from synchronized diurnal rhythms to stochastic, pulse-like bursts and high-frequency noise. We conduct an extensive evaluation of statistical, deep learning, and foundation models. Our experiments reveal a pivotal finding: while foundation models demonstrate superior zero-shot forecasting accuracy, this advantage does not inherently translate into better decision utility. Of practical significance, we systematically analyze how the selection of predictive quantiles acts as a critical lever. We provide actionable guidelines for calibrating these selections to balance the trade-off between resource efficiency and service reliability, offering vital insights for real-world deployment decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。