arXiv:2502.13171cs.CRcs.AI2025-02被引 5

无需配对比较,实时检测整套钓鱼攻击,保护隐私且抗生成式AI伪造。

Web Phishing Net (WPN): A scalable machine learning approach for real-time phishing campaign detection

  • 基于无监督学习,避免逐对比较,实现高可扩展性。
  • 一次检测整套钓鱼活动,对生成式AI伪造的链接仍保持高识别率。
  • 不分析用户内容,兼顾隐私保护,适合大规模实时监控场景。

网络钓鱼是当今最普遍的网络攻击类型,也是数据泄露的主要来源,对个人和企业造成重大影响。基于网页的钓鱼攻击最为常见,通过社交媒体或邮件中的链接诱导点击,导致系统暴露于更严重的攻击。现有检测方法多依赖有监督学习,需大量数据训练且计算开销大,还常分析邮件内容,侵犯用户隐私。此外,面对生成式AI技术催生的新型钓鱼链接,传统方法易被绕过。此前的无监督方法如聚类虽被尝试,但因成对比较导致可扩展性差,且检测率不高。本文提出一种新型无监督学习方法(Web Phishing Net, WPN),无需成对比较,具备高速与可扩展性,可一次性检测整个钓鱼攻击集群,对由恶意实体利用生成式AI生成的定向钓鱼链接也具有高检测率,同时保障用户隐私。

原文摘要 · Abstract (English)

Phishing is the most prevalent type of cyber-attack today and is recognized as the leading source of data breaches with significant consequences for both individuals and corporations. Web-based phishing attacks are the most frequent with vectors such as social media posts and emails containing links to phishing URLs that once clicked on render host systems vulnerable to more sinister attacks. Research efforts to detect phishing URLs have involved the use of supervised learning techniques that use large amounts of data to train models and have high computational requirements. They also involve analysis of features derived from vectors including email contents thus affecting user privacy. Additionally, they suffer from a lack of resilience against evolution of threats especially with the advent of generative AI techniques to bypass these systems as with AI-generated phishing URLs. Unsupervised methods such as clustering techniques have also been used in phishing detection in the past, however, they are at times unscalable due to the use of pair-wise comparisons. They also lack high detection rates while detecting phishing campaigns. In this paper, we propose an unsupervised learning approach that is not only fast but scalable, as it does not involve pair-wise comparisons. It is able to detect entire campaigns at a time with a high detection rate while preserving user privacy; this includes the recent surge of campaigns with targeted phishing URLs generated by malicious entities using generative AI techniques.

钓鱼检测无监督学习生成式AI对抗实时检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。