arXiv:2603.29537cs.CRcs.AI2026-03

用流混合机制提升加密流量分类的上下文理解能力

Mean Masked Autoencoder with Flow-Mixing for Encrypted Traffic Classification

  • 通过教师-学生架构与流混合策略,实现跨流上下文学习
  • 在多个数据集上达到当前最优性能,显著超越传统MAE方法
  • 适合需要高精度加密流量分析的网络安全研究者

基于掩码自编码器(MAE)的自监督预训练模型在网络流量分类中展现出巨大潜力。然而,现有方法仅局限于单一流的字节级重建,缺乏对流量多粒度上下文关系的充分感知。为此,本文提出均值掩码自编码器(MMAE),一种采用流混合策略的教师-学生MAE范式,用于构建加密流量预训练模型。MMAE通过自蒸馏机制实现师生交互,教师提供未掩码流级语义监督,推动学生从局部字节重建迈向多粒度理解。为打破单一流的信息瓶颈,引入动态流混合(FlowMix)策略,替代传统随机掩码,通过构造带干扰的跨流混合样本,迫使模型从扭曲标记中学习判别性表示。此外,设计了包重要性感知掩码预测器(PMP),结合注意力偏置机制,利用包级侧信道统计信息动态掩码高语义密度的标记。大量实验在涵盖加密应用、恶意软件和攻击流量的多个数据集上验证了MMAE的优越性能。

原文摘要 · Abstract (English)

Network traffic classification using self-supervised pre-training models based on Masked Autoencoders (MAE) has demonstrated a huge potential. However, existing methods are confined to isolated byte-level reconstruction of individual flows, lacking adequate perception of the multi-granularity contextual relationship in traffic. To address this limitation, we propose Mean MAE (MMAE), a teacher-student MAE paradigm with flow mixing strategy for building encrypted traffic pre-training model. MMAE employs a self-distillation mechanism for teacher-student interaction, where the teacher provides unmasked flow-level semantic supervision to advance the student from local byte reconstruction to multi-granularity comprehension. To break the information bottleneck in individual flows, we introduce a dynamic Flow Mixing (FlowMix) strategy to replace traditional random masking mechanism. By constructing challenging cross-flow mixed samples with interferences, it compels the model to learn discriminative representations from distorted tokens. Furthermore, we design a Packet-importance aware Mask Predictor (PMP) equipped with an attention bias mechanism that leverages packet-level side-channel statistics to dynamically mask tokens with high semantic density. Numerous experiments on a number of datasets covering encrypted applications, malware, and attack traffic demonstrate that MMAE achieves state-of-the-art performance. The code is available at https://github.com/lx6c78/MMAE

流量分类自监督学习加密分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。