用多模态引擎融合检测钓鱼邮件,准确率高且误报低。
A Hybrid, Multi-Layered Pipeline for Phishing and Threat Classification: Independently Validated URL and NLP Engines with a Calibrated Multi-Channel Fusion Stage
- 分模块用URL、NLP和情报同步引擎独立评估,再融合结果。
- 综合测试下F1达0.914,真实垃圾邮件误报仅3.6%。
- 适合需要高精度反钓鱼系统的安全团队部署使用。
钓鱼攻击具有多模态特性。本文提出一种混合式分层检测管道,对每种模态分别使用专用引擎打分并融合结果。构建并独立基准测试了三个引擎:四阶段URL分析栈(域名防护、词法模型、威胁情报与非对称L2融合辅助模块);经过泛化强化的DistilBERT NLP分类器,其在未见真实钓鱼样本上的召回率从0.8%提升至87.3%;以及具备端到端OpenTelemetry监控的威胁情报同步器,验证了消息1:1保真。在包含10,677封邮件的全系统基准上,决策级融合阶段采用校准的概率或门控策略,在URL、邮件头和钓鱼概率通道上实现F1=0.914,同时将未见真实垃圾邮件的误报率降至3.6%。由于该基准使用代理的URL与头信息通道,且当前运行点仍需重新校准,故视为初步集成结果。对于可部署检测而言,决定性能上限的是模型的泛化能力,而非在自身训练分布数据上的评分精度。
原文摘要 · Abstract (English)
Phishing is a multi-modal threat. We present a hybrid pipeline that scores each modality with its own engine and fuses the results. Three engines are built, deployed, and independently benchmarked: a four-stage URL stack (Domain Guard, lexical model, threat intelligence, and an asymmetric L2 fusion sidecar); a generalization-hardened DistilBERT NLP classifier whose held-out real-phishing recall rises from 0.8% to 87.3%; and a threat-intelligence synchronizer with end-to-end OpenTelemetry instrumentation confirming 1:1 message conservation. A decision-level fusion stage, characterized on a 10,677-email whole-system benchmark, reaches F1=0.914 with a calibrated probabilistic-OR over URL, header, and phishing-probability channels while reducing held-out real-spam false positives to 3.6%. Because that benchmark uses proxy URL and header channels and an operating point still needing recalibration, we present it as a preliminary integrated result. For deployable detection, the limiting factor is how well a model generalizes, not how accurately it scores data drawn from its own training distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。