发现并预警GPU悄然失效,靠结构信号而非数值异常
When GPUs Fail Quietly: Observability-Aware Early Warning Beyond Numeric Telemetry
- 结合温度漂移与监控链路退化信号进行联合建模
- 检测到的脱落型故障几乎无数值前兆,但结构信号明显崩溃
- 适合运维团队和大规模集群管理者使用
GPU节点在现代高性能计算和人工智能任务中至关重要,但许多故障不会立即表现为硬故障。部分不稳定性表现为缓慢的热或效率漂移,而另一类故障则突然发生且极少或没有数值先兆。这类脱落型故障在驱动或互连层面导致GPU不可用,主要可观测信号为结构性变化,包括设备指标消失和监控数据完整性下降。本文提出一种可观测性感知的早期预警框架,联合建模(i)利用利用率感知的热漂移特征,以及(ii)监控链路退化指标,如抓取延迟增加、样本丢失、时间序列断层及设备指标消失。该框架在德国哥廷根大学计算中心(GWDG)生产环境的GPU节点数据上评估,可关联GPU、节点、监控与调度信号。结果表明,脱落型故障几乎无数值前兆,主要通过结构化遥测崩溃体现;联合建模相比仅依赖GPU信号的检测,显著提升了预警提前期。研究使用的数据集已公开于https://doi.org/10.5281/zenodo.19052367。
原文摘要 · Abstract (English)
GPU nodes are central to modern HPC and AI workloads, yet many failures do not manifest as immediate hard faults. While some instabilities emerge gradually as weak thermal or efficiency drift, a significant class occurs abruptly with little or no numeric precursor. In these detachment-class failures, GPUs become unavailable at the driver or interconnect level and the dominant observable signal is structural, including disappearance of device metrics and degradation of monitoring payload integrity. This paper proposes an observability-aware early-warning framework that jointly models (i) utilization-aware thermal drift signatures in GPU telemetry and (ii) monitoring-pipeline degradation indicators such as scrape latency increase, sample loss, time-series gaps, and device-metric disappearance. The framework is evaluated on production telemetry from GPU nodes at GWDG, where GPU, node, monitoring, and scheduler signals can be correlated. Results show that detachment failures exhibit minimal numeric precursor and are primarily observable through structural telemetry collapse, while joint modeling increases early-warning lead time compared to GPU-only detection. The dataset used in this study is publicly available at https://doi.org/10.5281/zenodo.19052367.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。