arXiv:2509.07287cs.CRcs.AI2025-09被引 4

用触发标签机制让大模型自动生成可追踪的钓鱼邮件。

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm

  • 在普通大模型中嵌入触发-标签关联,生成带可检测标记的内容。
  • 四种场景下检测准确率超90%,且隐蔽性强、计算开销低。
  • 适合安全团队部署,用于识别由大模型生成的钓鱼邮件。

随着大语言模型的快速发展,其被恶意利用生成钓鱼内容的风险日益突出。攻击者可利用大模型生成无拼写错误、高度定制化且难以察觉的钓鱼邮件,现有语义检测方法难以有效识别。尽管部分基于大模型的检测方法已显现潜力,但存在计算成本高、依赖基础模型性能的问题,难以大规模应用。本文提出Paladin,通过多种插入策略将触发-标签关联嵌入通用大模型,构建可被检测的增强型模型。当生成钓鱼相关内容时,系统会自动注入可识别标签。我们设计了隐式与显式触发器及标签,涵盖四种典型场景。实验从隐蔽性、有效性与鲁棒性三方面评估,结果表明,本方法在所有场景下检测准确率均超过90%,显著优于基线方法。

原文摘要 · Abstract (English)

With the rapid development of large language models, the potential threat of their malicious use, particularly in generating phishing content, is becoming increasingly prevalent. Leveraging the capabilities of LLMs, malicious users can synthesize phishing emails that are free from spelling mistakes and other easily detectable features. Furthermore, such models can generate topic-specific phishing messages, tailoring content to the target domain and increasing the likelihood of success. Detecting such content remains a significant challenge, as LLM-generated phishing emails often lack clear or distinguishable linguistic features. As a result, most existing semantic-level detection approaches struggle to identify them reliably. While certain LLM-based detection methods have shown promise, they suffer from high computational costs and are constrained by the performance of the underlying language model, making them impractical for large-scale deployment. In this work, we aim to address this issue. We propose Paladin, which embeds trigger-tag associations into vanilla LLM using various insertion strategies, creating them into instrumented LLMs. When an instrumented LLM generates content related to phishing, it will automatically include detectable tags, enabling easier identification. Based on the design on implicit and explicit triggers and tags, we consider four distinct scenarios in our work. We evaluate our method from three key perspectives: stealthiness, effectiveness, and robustness, and compare it with existing baseline methods. Experimental results show that our method outperforms the baselines, achieving over 90% detection accuracy across all scenarios.

钓鱼检测大模型安全触发标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。