揭露加密流量分类中表示学习的虚假高准确率陷阱
The Sweet Danger of Sugar: Debunking Representation Learning for Encrypted Traffic Classification
- 发现模型性能依赖数据预处理中的虚假关联,而非真实特征
- 实测环境下模型准确率大幅下降,最高仅40%左右
- 提出Pcap-Encoder并呼吁建立更严谨的评估标准
近期大量研究受BERT等语言模型启发,利用表示学习模型构建网络流量表征,在加密流量分类任务中宣称达到高达98%的准确率。本文以网络专家视角重新审视这些结果,通过广泛分析发现:所谓高性能主要源于数据准备问题,使模型在微调阶段捕获到特征与标签间的虚假相关性,从而获得不切实际的高分。当此类捷径不存在时(如真实场景),模型表现显著下降。我们提出Pcap-Encoder,一种基于语言模型、专用于从协议头提取特征的表示学习模型,目前唯一能提供有效表征的模型。但其复杂性限制了实际应用。研究揭示数据集构建和训练流程中的缺陷,呼吁采用更科学的评估方法,并强调严格基准测试的重要性。
原文摘要 · Abstract (English)
Recently we have witnessed the explosion of proposals that, inspired by Language Models like BERT, exploit Representation Learning models to create traffic representations. All of them promise astonishing performance in encrypted traffic classification (up to 98% accuracy). In this paper, with a networking expert mindset, we critically reassess their performance. Through extensive analysis, we demonstrate that the reported successes are heavily influenced by data preparation problems, which allow these models to find easy shortcuts - spurious correlation between features and labels - during fine-tuning that unrealistically boost their performance. When such shortcuts are not present - as in real scenarios - these models perform poorly. We also introduce Pcap-Encoder, an LM-based representation learning model that we specifically design to extract features from protocol headers. Pcap-Encoder appears to be the only model that provides an instrumental representation for traffic classification. Yet, its complexity questions its applicability in practical settings. Our findings reveal flaws in dataset preparation and model training, calling for a better and more conscious test design. We propose a correct evaluation methodology and stress the need for rigorous benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。