arXiv:2411.01929cs.LGcs.AI2024-11被引 1

将网络流量转为文本生成,提升稀缺数据的合成质量。

Exploring the Landscape for Generative Sequence Models for Specialized Data Synthesis

  • 把数值型网络流量转化为文本,用语言模型生成数据。
  • 合成数据在质量与泛化性上超越现有先进方法。
  • 适合需要隐私保护数据的网络安全研究者使用。

人工智能研究常致力于开发能在复杂数据集上可靠泛化的模型,但在数据稀缺、复杂或难以获取的领域仍面临挑战。本文提出一种新方法,利用三种不同复杂度的生成模型,合成最具有挑战性的结构化数据之一:恶意网络流量。该方法独特地将数值数据转换为文本,将数据生成重构为语言建模任务,不仅增强了数据正则化,还显著提升了泛化能力与合成数据质量。大量统计分析表明,该方法在生成高保真合成数据方面优于当前最优生成模型。此外,我们系统研究了合成数据的应用、有效性及评估策略,为多领域中其作用提供了宝贵见解。代码与预训练模型已开源至Github,便于进一步探索与应用。关键词:数据合成,机器学习,流量生成,隐私保护数据,生成模型。

原文摘要 · Abstract (English)

Artificial Intelligence (AI) research often aims to develop models that can generalize reliably across complex datasets, yet this remains challenging in fields where data is scarce, intricate, or inaccessible. This paper introduces a novel approach that leverages three generative models of varying complexity to synthesize one of the most demanding structured datasets: Malicious Network Traffic. Our approach uniquely transforms numerical data into text, re-framing data generation as a language modeling task, which not only enhances data regularization but also significantly improves generalization and the quality of the synthetic data. Extensive statistical analyses demonstrate that our method surpasses state-of-the-art generative models in producing high-fidelity synthetic data. Additionally, we conduct a comprehensive study on synthetic data applications, effectiveness, and evaluation strategies, offering valuable insights into its role across various domains. Our code and pre-trained models are openly accessible at Github, enabling further exploration and application of our methodology. Index Terms: Data synthesis, machine learning, traffic generation, privacy preserving data, generative models.

数据合成生成模型网络安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。