arXiv:2510.22397cs.NIcs.LG2025-10被引 3

针对网络数据突发性波动,提出事件驱动的统一建模方法

NetBurst: Event-Centric Forecasting of Bursty, Intermittent Time Series

  • 将时间序列拆解为事件发生时刻与幅度,聚焦突发峰值而非连续数据
  • 在真实生产环境数据上,预测误差降低1.3至116倍,更准确捕捉极端事件分布
  • 适合运维人员快速定位异常模式,提升故障诊断与历史案例检索效率

网络运营商通过采集包数、字节速率或流体积等遥测数据监控基础设施,但要实现负载预测、异常诊断与历史事件检索等运维需求,仅靠原始数据不足。需构建紧凑的实体级表征,以捕捉每个实体时序数据的动态特征。现有时序基础模型多针对密集周期性数据(温和统计场景),而网络遥测数据处于野生场景:关键事件稀少,间隔不一的低活动期(潮退)与间歇性重尾极端值爆发(潮涌)并存。本文提出NetBurst,一种事件中心化流程,将每条时序数据分解为事件发生时刻流与事件幅度流,学习统一表示以支持预测、异常刻画与历史检索三类任务。在九种生产配置下,相较于八种基线模型(包括Amazon Chronos-2与Datadog Toto),NetBurst在野生场景中预测误差降低1.3–116倍,对真实突发分布的匹配度提升1.0–7.5倍;在温和场景中表现持平。异常刻画方面,生成的聚类更均衡且可解释性提升16倍(基于新设计的可解释性评分),结合聚类过滤的搜索实现7.5倍加速。

原文摘要 · Abstract (English)

Network operators monitor their infrastructure by collecting telemetry data such as packet counts, byte rates, or flow volumes, yet answering the questions that effective operations demand -- forecasting future load, diagnosing and characterizing anomalies, and searching for and retrieving historical precedents -- requires more than raw measurements. Bridging this gap calls for learned representations: compact per-entity summaries that capture temporal dynamics from each entity's univariate time series. Time-series foundation models are the natural starting point, but they are designed for dense, periodic benchmark datasets -- the \emph{mild} statistical regime. However, network telemetry data inhabits the \emph{wild} regime: operationally relevant events are rare, separated by variable-length stretches of low or no activity (``ebbs''), with intermittent bursts of heavy-tailed extremes (``tides''). We present NetBurst, an event-centric pipeline that collapses ebbs, separates each time series into a stream of burst timings and a stream of burst magnitudes, and learns a single representation serving all three operational tasks. Compared to the strongest competitors among eight baselines -- including Amazon's Chronos-2 and Datadog's Toto -- and across nine production telemetry configurations, NetBurst reduces median forecasting error by $1.3$--$116\times$ on wild-regime data with a $1.0$--$7.5\times$ better match to the true burst distribution, and matches baselines on mild-regime benchmarks. For characterizing anomalies, NetBurst produces balanced, well-spread clusters that are $16\times$ more describable in operator-familiar terms under a novel interpretability score, and cluster-filtered search delivers $7.5\times$ faster end-to-end retrieval.

时间序列异常检测事件建模运维智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。