arXiv:2507.17978cs.CRcs.AI2025-07被引 3

构建了含13万封邮件的多源钓鱼邮件数据集,提升检测模型性能。

MeAJOR Corpus: A Multi-Source Dataset for Phishing Email Detection

  • 整合13个开源库,覆盖多种钓鱼手法与正常邮件
  • 在XGB模型下达到98.34%的F1分数
  • 解决类别不平衡问题,适合安全研究与模型验证

钓鱼邮件持续通过欺骗性内容和恶意附件威胁网络安全,利用人类心理弱点。尽管机器学习模型在检测钓鱼威胁方面有效,但其表现高度依赖训练数据的质量与多样性。本文提出MeAJOR(Merged email Assets from Joint Open-source Repositories)数据集,一个新型多源钓鱼邮件数据集,旨在克服现有资源的关键局限。该数据集整合了135,894个样本,涵盖广泛的钓鱼策略与合法邮件,并包含多类工程特征。我们通过四类分类模型(RF、XGB、MLP、CNN)在多种特征配置下系统评估了数据集的实用性。结果表明,该数据集具有显著有效性,在XGB模型下实现98.34%的F1分数。通过融合多类别特征,本数据集提供可复用、一致的研究资源,有效应对类别不平衡、泛化能力与可重现性等常见挑战。

原文摘要 · Abstract (English)

Phishing emails continue to pose a significant threat to cybersecurity by exploiting human vulnerabilities through deceptive content and malicious payloads. While Machine Learning (ML) models are effective at detecting phishing threats, their performance largely relies on the quality and diversity of the training data. This paper presents MeAJOR (Merged email Assets from Joint Open-source Repositories) Corpus, a novel, multi-source phishing email dataset designed to overcome critical limitations in existing resources. It integrates 135894 samples representing a broad number of phishing tactics and legitimate emails, with a wide spectrum of engineered features. We evaluated the dataset's utility for phishing detection research through systematic experiments with four classification models (RF, XGB, MLP, and CNN) across multiple feature configurations. Results highlight the dataset's effectiveness, achieving 98.34% F1 with XGB. By integrating broad features from multiple categories, our dataset provides a reusable and consistent resource, while addressing common challenges like class imbalance, generalisability and reproducibility.

钓鱼检测数据集机器学习网络安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。