用强化学习模拟真实黑客,高效绕过安卓恶意软件检测
REPLICANT: Learning Policies for Evading and Hardening Malware Detectors

- 基于标签仅黑盒环境训练可复用的篡改策略
- 在7个检测器上平均攻击成功率78.8%,优于现有方法20.9%-39.2%
- 不仅提升攻击能力,还能用于训练更鲁棒的防御模型
为评估基于机器学习的恶意软件检测在现实世界中的有效性,必须考察其对强大对手的鲁棒性。然而,现有攻击方法难以模拟真实对手,常假设可访问训练数据、特征空间或置信度分数等特权信息。本文提出Replicant,一种深度强化学习框架,在严格的仅标签黑盒威胁模型下学习真实的规避任务。Replicant学习如何修改恶意软件样本及何时查询目标,该策略可在样本、检测器和特征空间间迁移。在七个Android恶意软件检测器和三个特征空间上,Replicant是表现最强且查询最高效的方案,平均攻击成功率达到78.8%,相对现有最佳方法提升20.9%至39.2%。此外,作为对抗训练工具,Replicant生成的检测器具备更强的泛化鲁棒性,优于当前最优水平。结果表明,学习规避任务不仅能增强攻击效果,更提供了更好的加固信号。
原文摘要 · Abstract (English)
To determine the real-world effectiveness of machine learning based malware detection, it is vital to evaluate its robustness against highly capable adversaries. However, state-of-the-art attacks do not effectively model realistic adversaries, as they often assume access to privileged information such as the training data, feature space, or confidence scores of the target. In this work, we present Replicant, a deep reinforcement learning framework that learns the realistic task of evasion under a strict label-only black-box threat model. Replicant learns a reusable policy on how to modify a malware sample and when to query the target, which transfers across samples, detectors, and feature spaces. Across seven Android malware detectors and three feature spaces, Replicant is the strongest and most query-efficient approach achieving a mean attack success rate of 78.8%, a relative improvement of 20.9%-39.2% over the state-of-the-art. Furthermore, when used for adversarial training, Replicant also outperforms the state-of-the art by producing detectors with more generalizable robustness. With Replicant we demonstrate that learning the task of evasion not only results in stronger attack performance but, crucially, provides a better signal for hardening malware detectors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。