构建综合评估框架,对比AI杀毒工具在真实场景下的表现优劣。
EXE-Bench: Ranking the Tradeoffs of AI-based Windows Malware Detectors for Real-World Usability

- 设计多维度评测基准,覆盖性能、时间稳定性与对抗攻击鲁棒性。
- 发现人工特征工程模型比深度网络更抗时间推移和恶意注入攻击。
- 适合安全研究人员与企业部署者参考,避免盲目采用新模型。
由于缺乏系统性评估,我们尚无法确定应部署哪种基于AI的Windows恶意软件检测器。现有评估存在四大缺陷:(i) 训练与测试数据不一致;(ii) 忽略时间维度分析,无法检验模型随时间退化情况;(iii) 未进行对抗攻击测试,难以暴露内容注入攻击下的脆弱性;(iv) 忽视部署时的计算开销,可能导致终端推理过慢。为此,我们提出EXE-Bench,一个全面评估基于AI的Windows恶意软件检测器的基准。该基准综合评估性能、时间鲁棒性、对抗鲁棒性及计算开销,并生成统一得分实现公平比较。通过分析发现,仅在部署后评估模型表现是不充分的,无法完整反映其真实能力。特别地,经过领域知识强化的特征工程方法在抵抗时间演化与对抗攻击方面表现优异,远超多数深度网络,后者仅在部署初期表现良好。
原文摘要 · Abstract (English)
Due to the lack of systematic evaluations, we are not yet able to determine which AI-based Windows malware detector to deploy in production, since existing evaluations (i) differ in terms of data used for both training and testing; (ii) do not consider temporal analysis to showcase whether models withstand the passage of time; (iii) avoid security evaluations with adversarial attacks that could highlight their brittleness against content-injection attacks; and (iv) neglect the computational requirements for deployment, risking slow inference on endpoints. For these reasons, we develop EXE-Bench, a comprehensive benchmark of AI-based Windows malware detectors. EXE-Bench assesses performance, temporal and adversarial robustness, and computational overhead, aggregating them into a single score for direct and fair model comparison. Through EXE-Bench, we highlight how evaluations conducted only after deployment are suboptimal and unable to provide a complete picture of their performance. In particular, through our analysis, we remark how much domain knowledge instilled through feature engineering is still extremely useful in this domain, resisting both time and adversarial attacks, in stark contrast with most of the deep networks that only excel right after deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。