arXiv:2509.01791cs.CRcs.AI2025-09中稿 · ACM AISec '25被引 11

提出新数据集生成框架,破解钓鱼邮件检测研究中的基准失真问题。

E-PhishGen: Unlocking Novel Research in Phishing Email Detection

  • 用大模型生成多语言钓鱼邮件,构建更贴近现实的新数据集
  • 新数据集上多数现有方法性能大幅下降,暴露检测能力瓶颈
  • 开源全部代码与数据,推动更真实、更具挑战性的研究

每日大量电子邮件涌入收件箱,从烦人的垃圾邮件到隐蔽的钓鱼诈骗。尽管已有研究宣称检测准确率接近完美,但实际对抗恶意邮件仍面临未解难题。本文对钓鱼邮件检测领域的基准数据集进行批判性评估,发现多数研究依赖不具代表性的英文数据集。为此,我们复现并重新评估多种机器学习与大语言模型方法,开源全部代码。结果表明,这些方法在原数据集上表现近乎完美,但迁移到新构建的E-PhishLLM数据集(含16616封多语言邮件)后性能显著下降。通过30人用户验证,E-PhishLLM具备高真实性。研究证明,钓鱼邮件检测仍是开放问题,并提供可复现的解决方案供后续研究使用。

原文摘要 · Abstract (English)

Every day, our inboxes are flooded with unsolicited emails, ranging between annoying spam to more subtle phishing scams. Unfortunately, despite abundant prior efforts proposing solutions achieving near-perfect accuracy, the reality is that countering malicious emails still remains an unsolved dilemma. This "open problem" paper carries out a critical assessment of scientific works in the context of phishing email detection. First, we focus on the benchmark datasets that have been used to assess the methods proposed in research. We find that most prior work relied on datasets containing emails that -- we argue -- are not representative of current trends, and mostly encompass the English language. Based on this finding, we then re-implement and re-assess a variety of detection methods reliant on machine learning (ML), including large-language models (LLM), and release all of our codebase -- an (unfortunately) uncommon practice in related research. We show that most such methods achieve near-perfect performance when trained and tested on the same dataset -- a result which intrinsically hinders development (how can future research outperform methods that are already near perfect?). To foster the creation of "more challenging benchmarks" that reflect current phishing trends, we propose E-PhishGEN, an LLM-based (and privacy-savvy) framework to generate novel phishing-email datasets. We use our E-PhishGEN to create E-PhishLLM, a novel phishing-email detection dataset containing 16616 emails in three languages. We use E-PhishLLM to test the detectors we considered, showing a much lower performance than that achieved on existing benchmarks -- indicating a larger room for improvement. We also validate the quality of E-PhishLLM with a user study (n=30). To sum up, we show that phishing email detection is still an open problem -- and provide the means to tackle such a problem by future research.

钓鱼邮件数据生成LLM应用安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。