用GPT生成数据中心流量数据,还原真实网络模式。
Harnessing Generative Pre-Trained Transformer for Datacenter Packet Trace Generation
- 基于GPT架构构建流量生成模型DTG-GPT,学习真实数据特征。
- 生成的新流量能复现原始数据的时空模式,支持不同规模网络。
- 适合研究者和运营商用于隐私保护下的数据共享与系统优化。
当前依赖数据中心的应用迅速增长,亟需新方法应对日益增加的流量与计算需求。数据中心流量数据对未来发展与优化至关重要,但公开数据极少。现有研究多采用简化数学模型,难以还原复杂流量模式,错失优化机会。本文提出基于生成式预训练变压器(GPT)架构的包级数据中心流量生成器DTG-GPT。在少量来自不同领域的可用流量数据上训练,并提供简单评估方法以衡量生成流量与原始数据的保真度。实验表明,DTG-GPT可生成模仿真实流量时空特性的新流量,且适用于不同规模网络。结果表明,未来类似模型或可使数据中心运营商通过发布训练好的GPT模型,安全地向研究社区共享流量信息。
原文摘要 · Abstract (English)
Today, the rapid growth of applications reliant on datacenters calls for new advancements to meet the increasing traffic and computational demands. Traffic traces from datacenters are essential for further development and optimization of future datacenters. However, traces are rarely released to the public. Researchers often use simplified mathematical models that lack the depth needed to recreate intricate traffic patterns and, thus, miss optimization opportunities found in realistic traffic. In this preliminary work, we introduce DTG-GPT, a packet-level Datacenter Traffic Generator (DTG), based on the generative pre-trained transformer (GPT) architecture used by many state-of-the-art large language models. We train our model on a small set of available traffic traces from different domains and offer a simple methodology to evaluate the fidelity of the generated traces to their original counterparts. We show that DTG-GPT can synthesize novel traces that mimic the spatiotemporal patterns found in real traffic traces. We further demonstrate that DTG-GPT can generate traces for networks of different scales while maintaining fidelity. Our findings indicate the potential that, in the future, similar models to DTG-GPT will allow datacenter operators to release traffic information to the research community via trained GPT models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。