arXiv:2602.15298cs.AI2026-02KDD

X-MAP通过语义模式分析,精准识别垃圾邮件与钓鱼攻击的误判样本。

X-MAP: eXplainable Misclassification Analysis and Profiling for Spam and Phishing Detection

  • 结合SHAP与非负矩阵分解,构建可解释的消息主题特征图谱
  • 误判样本的偏离度至少是正确分类样本的2倍,检测准确率达0.98 AUROC
  • 适合需要高可信度与可解释性的安全系统部署

垃圾邮件与钓鱼攻击检测中的误判危害极大:漏报导致用户暴露于攻击,误报则降低系统信任。现有基于不确定性的检测方法虽能标记潜在错误,但易被欺骗且解释性有限。本文提出X-MAP——一种可解释的误判分析与画像框架,揭示模型失败背后的主题级语义模式。X-MAP融合SHAP特征归因与非负矩阵分解,为正确分类的垃圾/钓鱼消息及合法消息构建可解释的主题画像,并使用Jensen-Shannon散度度量每条消息与画像的偏离程度。在SMS与钓鱼数据集上的实验表明,误判样本的偏离度至少是正确分类样本的2倍。作为检测器,X-MAP最高实现0.98 AUROC,于95%真阳性率下将误拒率降至0.089。作为修复层应用于基础检测器时,可恢复高达97%的误拒合法样本,且泄露可控。结果验证了X-MAP在提升垃圾邮件与钓鱼检测效果与可解释性方面的有效性。

原文摘要 · Abstract (English)

Misclassifications in spam and phishing detection are very harmful, as false negatives expose users to attacks while false positives degrade trust. Existing uncertainty-based detectors can flag potential errors, but possibly be deceived and offer limited interpretability. This paper presents X-MAP, an eXplainable Misclassification Analysis and Profilling framework that reveals topic-level semantic patterns behind model failures. X-MAP combines SHAP-based feature attributions with non-negative matrix factorization to build interpretable topic profiles for reliably classified spam/phishing and legitimate messages, and measures each message's deviation from these profiles using Jensen-Shannon divergence. Experiments on SMS and phishing datasets show that misclassified messages exhibit at least two times larger divergence than correctly classified ones. As a detector, X-MAP achieves up to 0.98 AUROC and lowers the false-rejection rate at 95% TRR to 0.089 on positive predictions. When used as a repair layer on base detectors, it recovers up to 97% of falsely rejected correct predictions with moderate leakage. These results demonstrate X-MAP's effectiveness and interpretability for improving spam and phishing detection.

可解释性垃圾邮件检测钓鱼攻击误判分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。