arXiv:2411.19493cs.NIcs.LG2024-11被引 10

用扩散模型提升网络流量矩阵估计精度,仅需5%数据即可达到良好效果。

Diffusion Models Meet Network Management: Improving Traffic Matrix Analysis with Diffusion-based Approach

  • 基于扩散模型生成真实流量分布,无需特定任务假设。
  • 在仅有5%已知值情况下仍能准确估计全量流量。
  • 适用于实际网络管理中数据缺失场景,即插即用且有理论保障。

由于网络运维高度依赖流量监测,流量矩阵分析成为网络管理的核心问题。然而,因测量成本高及传输损耗,精确获取网络流量极为困难。尽管近年已有方法从部分流级或链路级测量中估算流量,但在现代复杂动态网络中表现不佳。现有技术多依赖低秩结构或先验分布等强假设,且任务特定,泛化能力差。为此,本文提出基于扩散模型的流量矩阵分析框架Diffusion-TM,利用无任务依赖的扩散机制显著提升流量分布与精度估计性能。该框架不仅具备强大的生成能力以合成真实流量,还通过去噪过程实现端到端流量的无偏估计,支持即插即用并具理论保证。针对完整流量数据难以获取的问题,设计两阶段训练方案,使模型对数据缺失不敏感。大量真实数据集实验表明,Diffusion-TM在多项任务上表现优异,即使仅保留5%已知值也能获得理想结果。

原文摘要 · Abstract (English)

Due to network operation and maintenance relying heavily on network traffic monitoring, traffic matrix analysis has been one of the most crucial issues for network management related tasks. However, it is challenging to reliably obtain the precise measurement in computer networks because of the high measurement cost, and the unavoidable transmission loss. Although some methods proposed in recent years allowed estimating network traffic from partial flow-level or link-level measurements, they often perform poorly for traffic matrix estimation nowadays. Despite strong assumptions like low-rank structure and the prior distribution, existing techniques are usually task-specific and tend to be significantly worse as modern network communication is extremely complicated and dynamic. To address the dilemma, this paper proposed a diffusion-based traffic matrix analysis framework named Diffusion-TM, which leverages problem-agnostic diffusion to notably elevate the estimation performance in both traffic distribution and accuracy. The novel framework not only takes advantage of the powerful generative ability of diffusion models to produce realistic network traffic, but also leverages the denoising process to unbiasedly estimate all end-to-end traffic in a plug-and-play manner under theoretical guarantee. Moreover, taking into account that compiling an intact traffic dataset is usually infeasible, we also propose a two-stage training scheme to make our framework be insensitive to missing values in the dataset. With extensive experiments with real-world datasets, we illustrate the effectiveness of Diffusion-TM on several tasks. Moreover, the results also demonstrate that our method can obtain promising results even with $5\%$ known values left in the datasets.

流量矩阵扩散模型网络管理数据缺失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。