用Ksurf优化云资源分配,降低延迟波动和成本
Ksurf-Drone: Attention Kalman Filter for Contextual Bandit Optimization in Cloud Resource Allocation
- 引入Ksurf注意力卡尔曼滤波器改进带上下文的强化学习决策
- 延迟95分位降低41%,99分位降低47%,CPU与内存使用下降
- 适合高波动云环境下的资源调度优化,如Kubernetes集群
容器化云数据中心中的资源编排与参数搜索是关键挑战。面对庞大的配置空间和云环境不确定性,当前先进的Drone编排器采用上下文关联的多臂老虎机技术进行资源管理。然而,虚拟机数量变化带来的工作负载与资源指标波动,加剧了非线性与噪声,影响编排精度。本文将Ksurf——一种针对高度可变云数据的方差最小化估计方法——作为上下文多臂老虎机的目标函数模型,应用于基于Drone的资源估计任务。实验表明,在变异性高的工作负载下,Ksurf使延迟的p95分位降低41%,p99分位降低47%;在Kubernetes上实现CPU使用减少4%、主节点内存降低7MB;在VarBench基准测试中,平均工作节点数减少7%,带来7%的成本节约。
原文摘要 · Abstract (English)
Resource orchestration and configuration parameter search are key concerns for container-based infrastructure in cloud data centers. Large configuration search space and cloud uncertainties are often mitigated using contextual bandit techniques for resource orchestration including the state-of-the-art Drone orchestrator. Complexity in the cloud provider environment due to varying numbers of virtual machines introduces variability in workloads and resource metrics, making orchestration decisions less accurate due to increased nonlinearity and noise. Ksurf, a state-of-the-art variance-minimizing estimator method ideal for highly variable cloud data, enables optimal resource estimation under conditions of high cloud variability. This work evaluates the performance of Ksurf on estimation-based resource orchestration tasks involving highly variable workloads when employed as a contextual multi-armed bandit objective function model for cloud scenarios using Drone. Ksurf enables significantly lower latency variance of $41\%$ at p95 and $47\%$ at p99, demonstrates a $4\%$ reduction in CPU usage and 7 MB reduction in master node memory usage on Kubernetes, resulting in a $7\%$ cost savings in average worker pod count on VarBench Kubernetes benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。