arXiv:2508.16119cs.NIcs.AI2025-08

ANSC通过概率评估数据中⼼容量健康度,提前预警潜在拥堵风险。

ANSC: Probabilistic Capacity Health Scoring for Datacenter-Scale Reliability

  • 基于容量剩余与故障概率,动态计算健康评分
  • 覆盖400+数据中心、60+区域,减少误报噪音
  • 适合SRE团队用于识别最紧迫的可靠性风险

我们提出ANSC,一种用于超大规模数据中心网络的容量健康度概率评分框架。现有告警系统仅能检测单个设备或链路故障,却无法捕捉级联容量不足的累积风险。ANSC提供一种彩色评分系统,不仅依据当前影响,更根据即将发生容量违规的概率来判断问题紧急程度。该系统综合考虑当前剩余容量与额外故障概率,并在数据中心及区域层级进行归一化处理。实验表明,ANSC使运维人员能够在超过400个数据中心和60个区域中优先处理修复任务,显著降低干扰,使SRE团队聚焦于最关键的可靠性风险。

原文摘要 · Abstract (English)

We present ANSC, a probabilistic capacity health scoring framework for hyperscale datacenter fabrics. While existing alerting systems detect individual device or link failures, they do not capture the aggregate risk of cascading capacity shortfalls. ANSC provides a color-coded scoring system that indicates the urgency of issues \emph{not solely by current impact, but by the probability of imminent capacity violations}. Our system accounts for both current residual capacity and the probability of additional failures, normalized at datacenter and regional level. We demonstrate that ANSC enables operators to prioritize remediation across more than 400 datacenters and 60 regions, reducing noise and aligning SRE focus on the most critical risks.

数据中心可靠性概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。