用多模态数据精准定位超算网络故障节点与类型
ClusterRCA: An End-to-End Approach for Network Fault Localization and Classification for HPC System
- 融合拓扑连接的网卡对特征,分析异构网络数据
- 分类+图推理结合,故障定位准确率高
- 适合超算系统故障诊断,跨场景鲁棒性强
网络故障诊断对高性能计算(HPC)系统至关重要,但现有方法因数据异构性和精度不足难以直接应用。本文提出一种新框架ClusterRCA,通过提取拓扑连接的网络接口控制器(NIC)对特征,分析HPC系统中的多样化多模态数据,实现故障节点定位与故障类型判断。ClusterRCA结合基于分类器和图结构的方法:先由状态分类器输出构建故障图,再在图上执行定制化随机游走以定位根因。在某全球顶级超算设备厂商提供的数据集上实验表明,ClusterRCA在诊断HPC网络故障方面具有高准确率,并在不同应用场景下保持稳健性能。
原文摘要 · Abstract (English)
Network failure diagnosis is challenging yet critical for high-performance computing (HPC) systems. Existing methods cannot be directly applied to HPC scenarios due to data heterogeneity and lack of accuracy. This paper proposes a novel framework, called ClusterRCA, to localize culprit nodes and determine failure types by leveraging multimodal data. ClusterRCA extracts features from topologically connected network interface controller (NIC) pairs to analyze the diverse, multimodal data in HPC systems. To accurately localize culprit nodes and determine failure types, ClusterRCA combines classifier-based and graph-based approaches. A failure graph is constructed based on the output of the state classifier, and then it performs a customized random walk on the graph to localize the root cause. Experiments on datasets collected by a top-tier global HPC device vendor show ClusterRCA achieves high accuracy in diagnosing network failure for HPC systems. ClusterRCA also maintains robust performance across different application scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。