arXiv:2502.07492cs.CRcs.CV2025-02被引 5

对抗攻击下仍能精准识别恶意软件归属,提升网络安全威胁溯源能力

RoMA: Robust Malware Attribution via Byte-level Adversarial Training with Global Perturbations and Adversarial Consistency Regularization

  • 通过全局扰动与一致性正则化生成更强对抗样本,提升模型鲁棒性
  • 在PGD攻击下仍保持80%以上准确率,远超次优方法(不足35%)
  • 适用于需要高可靠性的高级持续性威胁溯源场景

将高级持续性威胁(APT)恶意软件归因于其所属团伙对威胁情报和网络安全至关重要。然而,攻击者常隐藏身份,使归因具有对抗性。现有基于机器学习的归因模型虽有效,但极易受对抗攻击影响。例如,最先进的字节级模型MalConv在投影梯度下降(PGD)攻击下准确率从90%以上骤降至2%以下。本文首次将基于梯度的对抗训练应用于恶意软件归因,发现其鲁棒性和训练效率均有待提升。为此,提出RoMA——一种新型单步对抗训练方法,结合全局扰动生成增强对抗样本,并引入对抗一致性正则化以提升表示质量与抗扰能力。构建了包含多样化样本和真实类别不平衡的新型APT恶意软件数据集AMG18用于评估。大量实验表明,RoMA在对抗鲁棒性上显著优于七种对比方法(如在PGD攻击下达到80%以上鲁棒准确率,超过次优方法两倍以上),训练效率也更高(速度超过次优方法两倍),同时在非对抗场景下保持优异标准准确率。

原文摘要 · Abstract (English)

Attributing APT (Advanced Persistent Threat) malware to their respective groups is crucial for threat intelligence and cybersecurity. However, APT adversaries often conceal their identities, rendering attribution inherently adversarial. Existing machine learning-based attribution models, while effective, remain highly vulnerable to adversarial attacks. For example, the state-of-the-art byte-level model MalConv sees its accuracy drop from over 90% to below 2% under PGD (projected gradient descent) attacks. Existing gradient-based adversarial training techniques for malware detection or image processing were applied to malware attribution in this study, revealing that both robustness and training efficiency require significant improvement. To address this, we propose RoMA, a novel single-step adversarial training approach that integrates global perturbations to generate enhanced adversarial samples and employs adversarial consistency regularization to improve representation quality and resilience. A novel APT malware dataset named AMG18, with diverse samples and realistic class imbalances, is introduced for evaluation. Extensive experiments show that RoMA significantly outperforms seven competing methods in both adversarial robustness (e.g., achieving over 80% robust accuracy-more than twice that of the next-best method under PGD attacks) and training efficiency (e.g., more than twice as fast as the second-best method in terms of accuracy), while maintaining superior standard accuracy in non-adversarial scenarios.

恶意软件归因对抗训练APT威胁鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。