提出无监督框架自动清理加密流量数据,提升分类效率。
Unsupervised Dataset Cleaning Framework for Encrypted Traffic Classification
- 无需人工干预,自动识别并剔除无效网络流。
- 相比人工清洗,准确率仅下降2%~2.5%。
- 适合需要高效预处理的智能网络分类场景。
流量分类技术被广泛应用于企业与运营商网络中。随着移动设备普及,应用层加密日益普遍,传统深度包检测(DPI)方法难以区分加密流量。为此,人工智能尤其是机器学习成为解决加密流量分类的有效方案。但任何基于机器学习的方法都依赖高质量的数据清洗,以剔除无关协议、背景活动、控制面消息和长连接会话等无效流量。现有清洗方法依赖人工逐包审查,成本高且耗时。本文提出一种无监督框架,可自动清洗加密移动流量。在真实数据集上的评估表明,该方法相较人工清洗仅导致2%~2.5%的分类准确率下降。结果证明,该方法为基于机器学习的加密流量分类提供了高效可靠的预处理手段。
原文摘要 · Abstract (English)
Traffic classification, a technique for assigning network flows to predefined categories, has been widely deployed in enterprise and carrier networks. With the massive adoption of mobile devices, encryption is increasingly used in mobile applications to address privacy concerns. Consequently, traditional methods such as Deep Packet Inspection (DPI) fail to distinguish encrypted traffic. To tackle this challenge, Artificial Intelligence (AI), in particular Machine Learning (ML), has emerged as a promising solution for encrypted traffic classification. A crucial prerequisite for any ML-based approach is traffic data cleaning, which removes flows that are not useful for training (e.g., irrelevant protocols, background activity, control-plane messages, and long-lived sessions). Existing cleaning solutions depend on manual inspection of every captured packet, making the process both costly and time-consuming. In this poster, we present an unsupervised framework that automatically cleans encrypted mobile traffic. Evaluation on real-world datasets shows that our framework incurs only a 2%~2.5% reduction in classification accuracy compared with manual cleaning. These results demonstrate that our method offers an efficient and effective preprocessing step for ML-based encrypted traffic classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。