arXiv:2608.01648cs.LG2026-08中稿 · the 7th Internatio…

分析超算硬件错误预测的边界,发现规律性错误可准预测

Evaluating Forecasting Techniques for Hardware Errors on a Large-scale HPC System

论文配图:Evaluating Forecasting Techniques for Hardware Errors on a Large-scale HPC System
图 1 · 摘自论文原文
  • 用七年超算日志测试经典与深度学习模型的预测能力
  • 规律性错误预测准确率高,突发性错误仍难捕捉
  • 为未来故障预测提供实证依据,适合系统运维研究者

高性能计算(HPC)系统的硬件错误日志提供了异常行为的早期信号,但利用现代预测方法有效预报这些错误仍面临挑战。本文基于Theta超算七年的生产日志,评估了经典统计模型与深度学习模型在预测硬件错误动态方面的表现。结果表明,预测效果高度依赖于错误序列的时间结构:规律性且结构稳定的错误可被准确建模,尤其是采用时间特征的LSTM和Transformer架构;而稀疏且爆发主导的错误仍难以预测。本研究未提出可部署的故障预测框架,而是提供了预测有效性适用场景的实证指导,并指出了提升预测精度的潜在方向。

原文摘要 · Abstract (English)

Hardware error logs in high-performance computing (HPC) systems provide early signals of abnormal behavior, yet there remain challenges in effectively forecasting these errors using modern predictive methods. This work investigates the boundaries of applying time series forecasting to HPC hardware error dynamics. We use seven years of production logs from the Theta supercomputer to evaluate the predictive efficacy of classical statistical and deep learning models. Our results show that forecasting effectiveness depends strongly on the temporal structure of the error series: regularly occurring and structurally stable errors can be modeled accurately, particularly by LSTM and Transformer architectures with temporal features, while sparse and burst-dominated errors remain difficult to predict. Rather than proposing a deployment-ready failure prediction framework, this study provides empirical guidance on when forecasting is effective and highlights potential directions for improving forecasting accuracy in HPC hardware error analysis.

故障预测时间序列超算系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。