通过聚类提升网络流量矩阵预测精度,显著降低误差和链路负载偏差。
Improving Internet Traffic Matrix Prediction via Time Series Clustering
- 按流量时序模式聚类,分组训练提升模型拟合能力。
- 在Abilene和GÉANT数据集上RMSE分别降低92%和75%。
- 适合需要高精度流量预测的网络优化场景。
我们提出一种新框架,利用时间序列聚类改进基于深度学习的互联网流量矩阵(TM)预测。由于流量矩阵中的流具有多样的时间行为,用单一模型训练所有流会降低预测准确率。为此,我们设计了源聚类和直方图聚类两种策略,在模型训练前将具有相似时间模式的流分组。聚类形成更同质的数据子集,使模型能更有效捕捉潜在模式并提升泛化能力,优于全局预测方法。相比现有方法,本方案在Abilene和GÉANT数据集上分别将均方根误差(RMSE)降低92%和75%。在路由场景中,聚类预测还分别降低了18%和21%的最大链路利用率(MLU)偏差,体现了聚类在流量矩阵用于网络优化时的实际价值。
原文摘要 · Abstract (English)
We present a novel framework that leverages time series clustering to improve internet traffic matrix (TM) prediction using deep learning (DL) models. Traffic flows within a TM often exhibit diverse temporal behaviors, which can hinder prediction accuracy when training a single model across all flows. To address this, we propose two clustering strategies, source clustering and histogram clustering, that group flows with similar temporal patterns prior to model training. Clustering creates more homogeneous data subsets, enabling models to capture underlying patterns more effectively and generalize better than global prediction approaches that fit a single model to the entire TM. Compared to existing TM prediction methods, our method reduces RMSE by up to 92\% for Abilene and 75\% for GÉANT. In routing scenarios, our clustered predictions also reduce maximum link utilization (MLU) bias by 18\% and 21\%, respectively, demonstrating the practical benefits of clustering when TMs are used for network optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。