arXiv:2606.29909cs.CV2026-06

提出可解释的多模态加密流量分类框架,让模型决策过程更透明。

Traffic-CBM: A Structurally Interpretable Multimodal Framework for Encrypted Traffic Classification

论文配图:Traffic-CBM: A Structurally Interpretable Multimodal Framework for Encrypted Traffic Classification
图 1 · 摘自论文原文
  • 将流量统计、时间序列和字节特征分层映射为显式概念
  • 在多个数据集上达到平衡性能,且解释性优于传统端到端模型
  • 适合需要理解模型决策依据的安全分析人员

加密流量分类已取得良好性能,但其决策过程难以解释。现有方法通常将流统计、包序列和字节级表示融合为黑箱隐含特征,无法明确哪些证据驱动预测。本文提出Traffic-CBM,一种结构可解释的多模态加密流量分类框架。不同于直接融合异构流量信号,Traffic-CBM将其组织为统一的分层概念空间:分组流统计映射为统计概念,专用时序编码器从分离特征子空间学习时间概念,字节级证据进一步分解为包级与跨包概念。该设计将异构流量证据转化为显式概念表示,使不同层级证据更易分析。在多个加密流量基准上评估表明,Traffic-CBM实现竞争性且均衡的分类性能,同时提供比传统端到端融合模型更清晰的结构化解释接口。进一步分析显示,学习的概念空间被主动用于预测,为多模态流量证据提供了更清晰的结构性解释。

原文摘要 · Abstract (English)

Encrypted traffic classification has achieved strong performance, but its decision process remains difficult to interpret. Existing methods usually combine flow statistics, packet sequences, and byte-level representations into opaque latent features, making it unclear which type of evidence actually drives the prediction. In this paper, we propose Traffic-CBM, a structurally interpretable multimodal framework for encrypted traffic classification. Instead of directly fusing heterogeneous traffic signals into a black-box representation, Traffic-CBM organizes them into a unified hierarchical concept space. These concepts are not manually annotated semantic attributes; rather, they are scalar evidence summaries constrained by predefined traffic evidence groups. More specifically, grouped flow statistics are mapped to statistical concepts, dedicated temporal encoders learn temporal concepts from disjoint feature subspaces, and byte-level evidence is further organized into packet-level and cross-packet concepts. This design turns heterogeneous traffic evidence into an explicit concept representation and makes different levels of traffic evidence easier to analyze. We evaluate Traffic-CBM on multiple encrypted traffic benchmarks. Results show that it achieves competitive and balanced classification performance while providing a clearer structural interpretation interface than conventional end-to-end fusion models. Further analyses suggest that the learned concept space is actively used in the prediction process and provides a clearer structural explanation of multimodal traffic evidence.

加密流量可解释性多模态分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。