arXiv:2511.04728cs.CRcs.AI2025-11被引 1

提出可信度校准框架,评估大模型钓鱼邮件检测的可靠性。

Trustworthiness Calibration Framework for Phishing Email Detection Using Large Language Models

  • 构建三维度评估框架,量化模型可信度。
  • GPT-4在可信度指标上表现最优,超越其他模型。
  • 适合安全系统部署前的可靠性验证,尤其关注真实场景鲁棒性。

钓鱼邮件持续威胁在线通信,利用人类信任并以逼真语言和自适应策略绕过自动化过滤。尽管大语言模型(如GPT-4和LLaMA-3-8B)在文本分类中表现优异,但其在安全系统中的部署需超越基准性能,评估其可靠性。为此,本文提出可信度校准框架(TCF),一种可复现的方法,从校准性、一致性与鲁棒性三个维度评估钓鱼检测器。这三个维度整合为一个有界指标——可信度校准指数(TCI),并辅以跨数据集稳定性(CDS)指标,量化模型在不同数据集上的可信度稳定性。在SecureMail 2025、Phishing Validation 2024、CSDMC2010、Enron-Spam和Nazario五个语料库上,使用DeBERTa-v3-base、LLaMA-3-8B和GPT-4进行实验,结果显示GPT-4总体可信度最佳,其次为LLaMA-3-8B,DeBERTa-v3-base最弱。统计分析表明,可靠性与原始准确率独立,强调了在实际部署中进行可信度感知评估的重要性。该框架为基于大模型的钓鱼邮件检测提供了透明且可复现的依赖性评估基础。

原文摘要 · Abstract (English)

Phishing emails continue to pose a persistent challenge to online communication, exploiting human trust and evading automated filters through realistic language and adaptive tactics. While large language models (LLMs) such as GPT-4 and LLaMA-3-8B achieve strong accuracy in text classification, their deployment in security systems requires assessing reliability beyond benchmark performance. To address this, this study introduces the Trustworthiness Calibration Framework (TCF), a reproducible methodology for evaluating phishing detectors across three dimensions: calibration, consistency, and robustness. These components are integrated into a bounded index, the Trustworthiness Calibration Index (TCI), and complemented by the Cross-Dataset Stability (CDS) metric that quantifies stability of trustworthiness across datasets. Experiments conducted on five corpora, such as SecureMail 2025, Phishing Validation 2024, CSDMC2010, Enron-Spam, and Nazario, using DeBERTa-v3-base, LLaMA-3-8B, and GPT-4 demonstrate that GPT-4 achieves the strongest overall trust profile, followed by LLaMA-3-8B and DeBERTa-v3-base. Statistical analysis confirms that reliability varies independently of raw accuracy, underscoring the importance of trust-aware evaluation for real-world deployment. The proposed framework establishes a transparent and reproducible foundation for assessing model dependability in LLM-based phishing detection.

钓鱼检测大模型可信度评估安全应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。