arXiv:2503.20796cs.CRcs.AI2025-03被引 40

用AI解释黑盒检测结果,让用户看清钓鱼攻击的判断依据。

EXPLICATE: Enhancing Phishing Detection through Explainable AI and LLM-Powered Interpretability

  • 结合LIME、SHAP与大模型生成双层可解释性分析。
  • 检测准确率达98.4%,解释内容与模型预测一致性达96.8%。
  • 适合安全人员和普通用户使用,提升对AI检测的信任。

复杂钓鱼攻击已成为主要网络安全威胁,日益普遍且难以防范。尽管机器学习在检测中展现潜力,但其决策过程如同“黑箱”,缺乏透明性,削弱用户信任并影响有效应对。本文提出EXPLICATE框架,采用三组件架构:基于域名特征的机器学习分类器、融合LIME与SHAP的双解释层以提供细粒度特征洞察,以及利用DeepSeek v3将技术解释转化为通俗自然语言。实验表明,该框架在所有指标上达到98.4%的准确率,与现有深度学习方法相当,同时具备更强可解释性。框架生成的解释准确性达94.2%,且与模型预测的一致性为96.8%。我们还开发了全功能图形界面应用及轻量级Chrome插件,验证其多场景部署可行性。研究证明,高精度检测与有意义解释可在安全应用中并行实现,弥合自动化AI与用户信任之间的关键鸿沟。

原文摘要 · Abstract (English)

Sophisticated phishing attacks have emerged as a major cybersecurity threat, becoming more common and difficult to prevent. Though machine learning techniques have shown promise in detecting phishing attacks, they function mainly as "black boxes" without revealing their decision-making rationale. This lack of transparency erodes the trust of users and diminishes their effective threat response. We present EXPLICATE: a framework that enhances phishing detection through a three-component architecture: an ML-based classifier using domain-specific features, a dual-explanation layer combining LIME and SHAP for complementary feature-level insights, and an LLM enhancement using DeepSeek v3 to translate technical explanations into accessible natural language. Our experiments show that EXPLICATE attains 98.4 % accuracy on all metrics, which is on par with existing deep learning techniques but has better explainability. High-quality explanations are generated by the framework with an accuracy of 94.2 % as well as a consistency of 96.8\% between the LLM output and model prediction. We create EXPLICATE as a fully usable GUI application and a light Chrome extension, showing its applicability in many deployment situations. The research shows that high detection performance can go hand-in-hand with meaningful explainability in security applications. Most important, it addresses the critical divide between automated AI and user trust in phishing detection systems.

可解释AI钓鱼检测大模型安全可信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。