通过多视角解释技术提升Linux系统故障预测的可信赖度。
Understanding Online Failure Prediction in Linux Through Complementary Multi-View Explainability

- 融合特征选择与时序、因果分析实现故障诊断
- 跨工作负载检测准确率达91%-94%,误报率低于1%
- 适合系统运维人员理解预测依据,尤其关注早期预警
准确的在线故障预测(OFP)在操作系统环境中已被证明可行,但仅靠预测不足以支持实际应用。缺乏诊断洞察会使运维人员难以信任告警或决定响应策略。即使预测精度高,也难判断模型是捕捉了真实的故障过程,还是仅利用了特定工作负载的噪声和偶然相关性。本文报告了在Linux系统上构建并评估可解释OFP管道的实际经验。结合基于共识的特征选择、时序起始分析、子系统级因果分析及互补诊断机制,实现故障解读。在严格跨工作负载条件下,冻结训练数据后仍实现91%-94%的检测率,且误报率低于1%。然而,故障模式诊断对工作负载变化更为敏感,部分诊断机制对特定故障类型效果有限。研究得出三大结论:(i) 检测比诊断更鲁棒地适应工作负载变化;(ii) 早期预警能力依赖故障类型,时间范围为38至215秒;(iii) 未见故障模式无法可靠由已知模式推断,LOMO评估下准确率为0%。这些结果表明,互补性解释机制有助于识别预测信号是否反映可迁移的故障结构,以及诊断泛化在工作负载变化下的失效边界。
原文摘要 · Abstract (English)
Accurate Online Failure Prediction (OFP) has been shown to be feasible in Operating Systems (OSs) settings, but prediction alone is not sufficient for practical adoption. Without diagnostic insight, operators have limited basis to trust alerts or decide how to respond. Moreover, even when predictive accuracy is high, it is often unclear whether models are capturing meaningful failure processes or merely exploiting workload-specific noise and incidental correlations in telemetry. This paper reports a practical experience building and evaluating an explainable OFP pipeline for Linux OSs. We combine consensus-based feature selection for detection with temporal onset analysis, subsystemlevel causal analysis, and complementary diagnostic mechanisms to support failure interpretation. Evaluated under strict crossworkload conditions with frozen training artifacts, it achieved 91-94% detection on unseen workloads without retraining, while maintaining false alarm rates below 1%. However, failure mode diagnosis proved substantially more sensitive to workload shift, and several diagnostics mechanisms showed limited effectiveness for specific failure types. Our experience highlights three main lessons: i) detection generalizes more robustly than diagnosis across workload changes; ii) early-warning capability depends strongly on the failure mode, ranging from 38 to 215 seconds in our study; and iii) unseen failure modes are not reliably diagnosable from related training modes alone, providing 0% accuracy under Leave-One-Mode-Out (LOMO) evaluation. Taken together, these results show the value of complementary explainability mechanisms for interpreting accurate failure predictions, revealing when predictive signals reflect transferable failure structure and when diagnostic generalization breaks down under workload variation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。