arXiv:2602.02641cs.CRcs.AI2026-02中稿 · NeurIPS

用大模型零样本检测钓鱼网址,应对日益猖獗的智能诈骗威胁。

Benchmarking Large Language Models for Zero-shot and Few-shot Phishing URL Detection

  • 采用统一提示框架评估大模型在零样本和少样本下的钓鱼网址识别能力。
  • 少样本提示使多款大模型性能提升,召回率与准确率显著改善。
  • 适合安全研究者与防御系统开发者,快速部署低依赖检测方案。

统一资源定位符(URL)诞生于以连接为核心的早期,虽经HTTPS等被动防护手段增强,仍缺乏面向未来的安全、信任与抗欺诈机制。当前在人工智能主导的攻击环境中,网络犯罪分子利用生成式AI制造高度拟真的钓鱼网站与链接,已达到用户和传统工具难以区分的程度。尽管2024年生成式钓鱼仅占过滤绕过攻击的一小部分,但整体钓鱼攻击量自2022年以来激增超4000%,近半数攻击可逃避检测。面对威胁演化速度远超标注数据生成速度的现状,零样本与少样本学习结合大语言模型(LLM)成为兼具时效性与适应性的解决方案。本文构建统一的零样本与少样本提示基准体系,评估多款主流模型在平衡数据集上的表现,通过准确率、精确率、召回率、F1分数、AUROC和AUPRC等指标量化分类质量与实际检测效用。实验表明,少样本提示可普遍提升各类模型性能,揭示关键权衡关系。

原文摘要 · Abstract (English)

The Uniform Resource Locator (URL), introduced in a connectivity-first era to define access and locate resources, remains historically limited, lacking future-proof mechanisms for security, trust, or resilience against fraud and abuse, despite the introduction of reactive protections like HTTPS during the cybersecurity era. In the current AI-first threatscape, deceptive URLs have reached unprecedented sophistication due to the widespread use of generative AI by cybercriminals and the AI-vs-AI arms race to produce context-aware phishing websites and URLs that are virtually indistinguishable to both users and traditional detection tools. Although AI-generated phishing accounted for a small fraction of filter-bypassing attacks in 2024, phishing volume has escalated over 4,000% since 2022, with nearly 50% more attacks evading detection. At the rate the threatscape is escalating, and phishing tactics are emerging faster than labeled data can be produced, zero-shot and few-shot learning with large language models (LLMs) offers a timely and adaptable solution, enabling generalization with minimal supervision. Given the critical importance of phishing URL detection in large-scale cybersecurity defense systems, we present a comprehensive benchmark of LLMs under a unified zero-shot and few-shot prompting framework and reveal operational trade-offs. Our evaluation uses a balanced dataset with consistent prompts, offering detailed analysis of performance, generalization, and model efficacy, quantified by accuracy, precision, recall, F1 score, AUROC, and AUPRC, to reflect both classification quality and practical utility in threat detection settings. We conclude few-shot prompting improves performance across multiple LLMs.

钓鱼检测大模型零样本安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。