用两阶段训练提升低资源环境下攻击技术分类准确率
Cyber-Attack Technique Classification Using Two-Stage Trained Large Language Models
- 利用同标签辅助数据增强训练,分两阶段优化模型
- 在TRAM数据集上宏平均F1提升5-9个百分点
- 适合安全分析、威胁情报自动化场景
理解攻击模式对把握攻击者行为和采取有效防护措施至关重要。然而,多数新攻击信息以非结构化文本形式存在,给安全分析师收集信息带来挑战。本文提出一种句子分类系统,可识别来自网络安全威胁情报报告中自然语言描述的攻击技术。我们提出一种新方法,利用具有相同标签的辅助数据来提升低资源场景下的攻击分类性能。系统首先用增强数据训练模型,再仅用主数据进一步微调。通过TRAM数据集和MITRE ATT&CK框架验证,实验表明该方法在TRAM数据集上使宏平均F1提升5至9个百分点,同时保持微平均F1竞争力,优于基线模型。
原文摘要 · Abstract (English)
Understanding the attack patterns associated with a cyberattack is crucial for comprehending the attacker's behaviors and implementing the right mitigation measures. However, majority of the information regarding new attacks is typically presented in unstructured text, posing significant challenges for security analysts in collecting necessary information. In this paper, we present a sentence classification system that can identify the attack techniques described in natural language sentences from cyber threat intelligence (CTI) reports. We propose a new method for utilizing auxiliary data with the same labels to improve classification for the low-resource cyberattack classification task. The system first trains the model using the augmented training data and then trains more using only the primary data. We validate our model using the TRAM data1 and the MITRE ATT&CK framework. Experiments show that our method enhances Macro-F1 by 5 to 9 percentage points and keeps Micro-F1 scores competitive when compared to the baseline performance on the TRAM dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。