arXiv:2512.10104cs.CRcs.AI2025-12被引 2

用大模型检测钓鱼邮件,能达90%以上准确率

Phishing Email Detection Using Large Language Models

  • 设计多向量检测框架,应对提示注入等攻击
  • 实测三款前沿大模型检测准确率超90%
  • 揭示大模型检测系统仍存被攻破风险,适合安全研究者

钓鱼邮件是当前最普遍且影响深远的网络入侵手段之一。随着大语言模型(LLMs)在各类系统中广泛应用,其固有架构漏洞也面临日益复杂的钓鱼攻击威胁。现有大模型在部署于邮件安全系统前需大量加固,尤其针对联合多向量攻击。本文提出一种基于大模型的钓鱼邮件检测框架 LLMPEA,可识别提示注入、文本优化及多语言攻击等多种攻击方式。我们评估了 GPT-4o、Claude Sonnet 4 与 Grok-3 三款前沿大模型,并结合多种提示工程策略,分析其在真实场景下的可行性、鲁棒性与局限性。实验表明,大模型在钓鱼邮件检测中可达超过 90% 的准确率;但同时发现,基于大模型的检测系统仍可能被对抗攻击、提示注入和多语言攻击所利用。研究结果为实际部署中的大模型钓鱼检测提供了关键洞见。

原文摘要 · Abstract (English)

Email phishing is one of the most prevalent and globally consequential vectors of cyber intrusion. As systems increasingly deploy Large Language Models (LLMs) applications, these systems face evolving phishing email threats that exploit their fundamental architectures. Current LLMs require substantial hardening before deployment in email security systems, particularly against coordinated multi-vector attacks that exploit architectural vulnerabilities. This paper proposes LLMPEA, an LLM-based framework to detect phishing email attacks across multiple attack vectors, including prompt injection, text refinement, and multilingual attacks. We evaluate three frontier LLMs (e.g., GPT-4o, Claude Sonnet 4, and Grok-3) and comprehensive prompting design to assess their feasibility, robustness, and limitations against phishing email attacks. Our empirical analysis reveals that LLMs can detect the phishing email over 90% accuracy while we also highlight that LLM-based phishing email detection systems could be exploited by adversarial attack, prompt injection, and multilingual attacks. Our findings provide critical insights for LLM-based phishing detection in real-world settings where attackers exploit multiple vulnerabilities in combination.

钓鱼邮件大模型安全检测框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。