arXiv:2511.21448cs.CRcs.AI2025-11被引 4

用真实邮件生成2.3万封带元数据的钓鱼/垃圾/正常邮件,用于评测大模型安全检测能力

The Phish, The Spam, and The Valid: Generating Feature-Rich Emails for Benchmarking LLMs

  • 将真实邮件注入大模型,控制实体与长度生成多样变体
  • 在23,100封邮件上测试两模型,发现结构元数据显著提升识别准确率
  • 开源生成框架与数据集,支持开放科学评测下一代邮件安全系统

本文提出一种元数据丰富的生成框架PhishFuzzer,将真实邮件输入大语言模型(LLMs),生成23,100个结构一致、多样化且可控实体与长度的邮件变体。与以往数据集不同,本数据集具备严格的三类标签(钓鱼、垃圾、有效)、完整的URL与附件元数据,并标注每封邮件的攻击者意图。我们基于该数据集,在Basic(仅正文和主题)与Full(含URL、发件人、附件)设置下,对两个前沿大模型(Qwen-2.5-72B和Gemini-3.1-Pro)进行基准测试。通过任务成功率与置信度指数等正式指标,分析模型可靠性、对抗语言模糊的能力,以及结构元数据对检测精度的影响。所提全开源框架与数据集为下一代邮件安全系统评估提供了严谨基础。相关代码、提示词与数据集已公开于GitHub:https://github.com/DataPhish/PhishFuzzer

原文摘要 · Abstract (English)

In this paper, we introduce a metadata-enriched generation framework (PhishFuzzer) that seeds real emails into Large Language Models (LLMs) to produce 23,100 diverse, structurally consistent email variants across controlled entity and length dimensions. Unlike prior corpora, our dataset features strict three-class labels (Phishing, Spam, Valid), provides full URL and attachment metadata, and annotates each email with attacker intent. Using this dataset, we benchmark two state-of-the-art LLMs (Qwen-2.5-72B and Gemini-3.1-Pro) under both Basic (body, subject) and Full (+URL, sender, attachment) settings. By applying formal confidence metrics (Task Success Rate and Confidence Index), we analyze model reliability, robustness against linguistic fuzzing, and the impact of structural metadata on detection accuracy. Our fully open-source framework and dataset provide a rigorous foundation for evaluating next-generation email security systems. To support open science, we make the PhishFuzzer Dataset, the generation scripts and prompts available on GitHub: https://github.com/DataPhish/PhishFuzzer

邮件安全大模型评测生成数据元数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。