通过因果序列分析提前定位基站网络故障根源。
Causal Intervention Sequence Analysis for Fault Tracking in Radio Access Networks

- 基于历史故障数据标注,学习从正常到故障的因果演变路径。
- 蒙特卡洛测试中准确识别故障触发顺序,支持百万级数据实时处理。
- 适合网络运维人员用于故障预防而非事后排查。
为保障现代无线接入网(RAN)稳定运行,运营商需在用户感知前识别服务等级协议(SLA)违约的真实诱因。本文提出一种AI/ML流程,弥补现有工具两大缺失:(1)定位可能的根本原因指标,(2)揭示事件发生的精确时序。通过标注历史数据——与过去SLA违约相关的记录标记为'异常',其余为'正常',模型学习从正常状态演变为故障的因果链条。蒙特卡洛测试表明,该方法以高精度定位正确触发序列,且在处理百万级数据点时仍保持高效。结果表明,高分辨率、因果有序的洞察可将故障管理从被动排障转向主动预防。
原文摘要 · Abstract (English)
To keep modern Radio Access Networks (RAN) running smoothly, operators need to spot the real-world triggers behind Service-Level Agreement (SLA) breaches well before customers feel them. We introduce an AI/ML pipeline that does two things most tools miss: (1) finds the likely root-cause indicators and (2) reveals the exact order in which those events unfold. We start by labeling network data: records linked to past SLA breaches are marked `abnormal', and everything else `normal'. Our model then learns the causal chain that turns normal behavior into a fault. In Monte Carlo tests the approach pinpoints the correct trigger sequence with high precision and scales to millions of data points without loss of speed. These results show that high-resolution, causally ordered insights can move fault management from reactive troubleshooting to proactive prevention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。