用通用流量表示增强大模型,让其在多种网络流量任务中表现更优。
TrafficLLM: Enhancing Large Language Models for Network Traffic Analysis with Generic Traffic Representation
- 通过双阶段微调学习异构原始流量的通用表征。
- 在229类流量上检测与生成任务中分别达0.9875和0.9483的F1分数。
- 对未见过的流量有18.6%性能提升,适合企业级实际部署。
基于机器学习的网络流量分析广泛用于威胁检测,但跨任务和未见数据的泛化能力较差。大语言模型(LLMs)虽具强大泛化能力,却因网络流量特性差异难以应用。为此,本文提出TrafficLLM,采用双阶段微调框架,从异构原始流量数据中学习通用流量表征。该框架结合流量领域分词、双阶段调优流程与可扩展适配机制,使LLM在动态流量分析任务中释放泛化潜力,实现多种下游任务中的流量检测与生成。我们在10种不同场景、229类流量上评估TrafficLLM,检测与生成任务的F1分数分别达到0.9875和0.9483,较现有方法提升最高达80.12%和33.92%。在未见过的流量上,性能提升18.6%。真实场景测试表明,TrafficLLM易于扩展,可在企业流量中实现精准检测。
原文摘要 · Abstract (English)
Machine learning (ML) powered network traffic analysis has been widely used for the purpose of threat detection. Unfortunately, their generalization across different tasks and unseen data is very limited. Large language models (LLMs), known for their strong generalization capabilities, have shown promising performance in various domains. However, their application to the traffic analysis domain is limited due to significantly different characteristics of network traffic. To address the issue, in this paper, we propose TrafficLLM, which introduces a dual-stage fine-tuning framework to learn generic traffic representation from heterogeneous raw traffic data. The framework uses traffic-domain tokenization, dual-stage tuning pipeline, and extensible adaptation to help LLM release generalization ability on dynamic traffic analysis tasks, such that it enables traffic detection and traffic generation across a wide range of downstream tasks. We evaluate TrafficLLM across 10 distinct scenarios and 229 types of traffic. TrafficLLM achieves F1-scores of 0.9875 and 0.9483, with up to 80.12% and 33.92% better performance than existing detection and generation methods. It also shows strong generalization on unseen traffic with an 18.6% performance improvement. We further evaluate TrafficLLM in real-world scenarios. The results confirm that TrafficLLM is easy to scale and achieves accurate detection performance on enterprise traffic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。