arXiv:2505.00034cs.CLcs.AI2025-05被引 3

用小模型提升钓鱼邮件检测,兼顾效果与算力

Improving Phishing Email Detection Performance of Small Large Language Models

  • 通过提示工程、解释增强微调和模型集成改进小模型性能
  • 在两个数据集上准确率和F1分数显著提升,接近大模型表现
  • 模型具备强迁移能力,适合资源受限场景下的安全应用

大型语言模型(LLMs)在自然语言处理任务中表现出色,并被用于钓鱼邮件检测研究。然而,现有研究中的高性能模型通常包含数十亿甚至上百亿参数,需要巨大计算资源。为降低计算成本,我们研究了参数量约30亿的小型语言模型在钓鱼邮件检测中的有效性,这类模型可在消费级显卡上运行。但小模型在该任务上表现不佳。为此,我们设计了一套方法,包括提示工程、解释增强微调和模型集成,以提升小模型的检测能力。实验验证了该方法的有效性,在SpamAssassin和CEAS_08数据集上显著提升了准确率和F1分数。此外,微调后的模型展现出强泛化能力,在多个未见过的钓鱼邮件数据集上表现稳健,优于传统基线方法,并接近标准尺寸大模型的表现。

原文摘要 · Abstract (English)

Large language models(LLMs) have demonstrated remarkable performance on many natural language processing(NLP) tasks and have been employed in phishing email detection research. However, in current studies, well-performing LLMs typically contain billions or even tens of billions of parameters, requiring enormous computational resources. To reduce computational costs, we investigated the effectiveness of small-parameter LLMs for phishing email detection. These LLMs have around 3 billion parameters and can run on consumer-grade GPUs. However, small LLMs often perform poorly in phishing email detection task. To address these issues, we designed a set of methods including Prompt Engineering, Explanation Augmented Fine-tuning, and Model Ensemble to improve phishing email detection capabilities of small LLMs. We validated the effectiveness of our approach through experiments, significantly improving both accuracy and F1 score on the SpamAssassin and CEAS\_08 datasets. Furthermore, the fine-tuned models demonstrated strong transferability, achieving robust performance across multiple unseen phishing datasets, outperforming traditional baselines and approaching standard-sized LLMs.

钓鱼邮件检测小模型提示工程模型集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。