通过两阶段在线学习提升移动核心网故障检测精度
Adaptive Two-Stage Online Learning for Service-Affecting Failure Detection in Mobile Core Networks

- 分两阶段建模:先拟合正常流量,再分析残差与上下文信号
- 在真实数据上实现高召回率与高F1值,误报率可控
- 适合需要实时监测的网络运维团队使用
移动网络运营商通过监控聚合流量来评估核心网络基础设施的运行健康状态。由于存在强时间结构、非平稳性、测量误差及极端类别不平衡,基于静态阈值的监测方法难以可靠地进行故障检测。本文提出一种两阶段在线学习框架,用于基于流量的移动核心网故障检测。第一阶段采用带时间感知特征的轻量级回归模型,逐步建模正常流量动态;第二阶段结合预测残差与上下文指标,识别真正影响服务的网络故障。该框架在预序评估协议下全在线运行,支持持续自适应且计算开销低。在线性与非线性模型中,所提两阶段架构均取得最优精确率-召回率权衡,达到最高召回率、F1分数和AUC,同时保持可接受的误报率。结果表明,在流式移动核心网数据中,显式残差分解对可靠故障检测至关重要。
原文摘要 · Abstract (English)
Mobile network operators monitor aggregated traffic volumes to assess the operational health of core network infrastructure. Reliable failure detection is challenging due to strong temporal structure, non-stationarity, measurement artefacts, and extreme class imbalance, which limit static threshold-based monitoring. This paper proposes a two-stage online learning framework for traffic-based failure detection in mobile core networks. Stage I incrementally models normal traffic dynamics using lightweight regression with time-aware features. Stage II analyses prediction residuals together with contextual indicators to detect genuine service-affecting network failures. The framework operates fully online under a prequential evaluation protocol, enabling continuous adaptation with low computational overhead. Across linear and non-linear models, the proposed two-stage architecture achieves the best precision-recall trade-off, attaining the highest recall, F1-score, and AUC at acceptable false positive rates. These results demonstrate the importance of explicit residual decomposition for reliable failure detection in streaming mobile core network data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。