用带领域约束的自编码器从特征向量重建网络会话数据
Reconstructing Fine-Grained Network Data using Autoencoder Architectures with Domain Knowledge Penalties
- 在自编码器中加入领域知识约束,提升重建精度
- 对分类特征的会话级编码,重建准确率显著提高
- 适合隐私敏感场景下的数据高效训练
从粗粒度特征向量重构细粒度网络会话数据(包括单个数据包)对提升网络安全模型至关重要。然而,大规模原始网络流量的采集与存储带来显著挑战,尤其在捕捉罕见网络攻击样本方面。这限制了全面数据集的构建,影响模型训练和未来威胁检测。为此,我们提出一种受形式化方法引导的机器学习方法,用于编码与重构网络数据。该方法采用带有领域信息惩罚项的自编码器,从结构化特征表示中恢复PCAP会话头信息。实验表明,通过基于约束的损失项引入领域知识,显著提升了重建准确性,特别是在会话级编码的分类特征上。本方法实现了详细网络会话的高效重建,支持数据高效的模型训练,同时保障隐私与存储效率。
原文摘要 · Abstract (English)
The ability to reconstruct fine-grained network session data, including individual packets, from coarse-grained feature vectors is crucial for improving network security models. However, the large-scale collection and storage of raw network traffic pose significant challenges, particularly for capturing rare cyberattack samples. These challenges hinder the ability to retain comprehensive datasets for model training and future threat detection. To address this, we propose a machine learning approach guided by formal methods to encode and reconstruct network data. Our method employs autoencoder models with domain-informed penalties to impute PCAP session headers from structured feature representations. Experimental results demonstrate that incorporating domain knowledge through constraint-based loss terms significantly improves reconstruction accuracy, particularly for categorical features with session-level encodings. By enabling efficient reconstruction of detailed network sessions, our approach facilitates data-efficient model training while preserving privacy and storage efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。