arXiv:2604.08140cs.CRcs.AI2026-04

用多模态大模型解析加密流量,生成可审计的解释报告。

Multimodal Reasoning with LLM for Encrypted Traffic Interpretation: A Benchmark

  • 构建字节级标注数据集BGTD,融合原始流量与专家语义标注。
  • 提出mmTraffic框架,实现高精度分类与可读性解释报告生成。
  • 适合安全分析、网络推理研究者,推动透明化流量理解。

网络流量作为现代互联网基础设施的关键媒介,对保障安全与通信至关重要。现有方法虽表现优异,但面临两大瓶颈:(1) 无法捕捉超越单模态序列模式的多维语义;(2) 黑箱特性仅输出类别标签,缺乏可审计的推理过程。我们发现,现有流量数据集主要面向分类任务,缺乏丰富语义标注,难以生成人类可读的证据报告。为此,本文首次提出字节级流量描述基准BGTD,结合原始字节与结构化专家标注,提供行为特征与可验证的证据链,支持可解释的加密流量多模态推理。基于BGTD,本文提出端到端流量-语言表征框架mmTraffic,一种连接物理流量编码与语义解释的多模态推理架构。为缓解模态干扰与生成幻觉,mmTraffic采用感知-认知联合优化架构,通过以感知为中心的流量编码器与以认知为中心的LLM生成器,实现精准且可解释的流量解析。大量实验表明,mmTraffic能自主生成高保真、可读、证据确凿的流量解释报告,同时在分类准确率上媲美专用单模态模型(如NetMamba)。源代码已公开于https://github.com/lgzhangzlg/Multimodal-Reasoning-with-LLM-for-Encrypted-Traffic-Interpretation-A-Benchmark。

原文摘要 · Abstract (English)

Network traffic, as a key media format, is crucial for ensuring security and communications in modern internet infrastructure. While existing methods offer excellent performance, they face two key bottlenecks: (1) They fail to capture multidimensional semantics beyond unimodal sequence patterns. (2) Their black box property, i.e., providing only category labels, lacks an auditable reasoning process. We identify a key factor that existing network traffic datasets are primarily designed for classification and inherently lack rich semantic annotations, failing to generate human-readable evidence report. To address data scarcity, this paper proposes a Byte-Grounded Traffic Description (BGTD) benchmark for the first time, combining raw bytes with structured expert annotations. BGTD provides necessary behavioral features and verifiable chains of evidence for multimodal reasoning towards explainable encrypted traffic interpretation. Built upon BGTD, this paper proposes an end-to-end traffic-language representation framework (mmTraffic), a multimodal reasoning architecture bridging physical traffic encoding and semantic interpretation. In order to alleviate modality interference and generative hallucinations, mmTraffic adopts a jointly-optimized perception-cognition architecture. By incorporating a perception-centered traffic encoder and a cognition-centered LLM generator, mmTraffic achieves refined traffic interpretation with guaranteed category prediction. Extensive experiments demonstrate that mmTraffic autonomously generates high-fidelity, human-readable, and evidence-grounded traffic interpretation reports, while maintaining highly competitive classification accuracy comparing to specialized unimodal model (e.g., NetMamba). The source code is available at https://github.com/lgzhangzlg/Multimodal-Reasoning-with-LLM-for-Encrypted-Traffic-Interpretation-A-Benchmark

加密流量分析多模态推理可解释AILLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。