arXiv:2506.21106cs.CRcs.AI2025-06

用关键词聚类提升钓鱼网页检测准确率,抗干扰能力强。

PhishKey: A Novel Centroid-Based Approach for Enhanced Phishing Detection Using Adaptive HTML Component Extraction

  • 通过聚类提取网页关键组件,自动去噪并保留完整内容。
  • 在四个数据集上最高达98.70%的F1分数,抗注入攻击性能好。
  • 适合需要高精度、强鲁棒性的网络安全防护场景。

钓鱼攻击构成重大网络安全威胁,持续演进以绕过检测机制并利用人类弱点。本文提出PhishKey,一种新型钓鱼检测方法,通过从混合源自动提取特征实现更强的适应性、鲁棒性和效率。PhishKey结合字符级处理与卷积神经网络(CNN)进行URL分类,并采用基于质心的关键组件钓鱼提取器(CAPE)在词级别处理HTML内容。CAPE能有效降低噪声,确保样本完整处理,避免输入数据裁剪操作。两个模块的预测结果通过软投票集成,实现更准确可靠的分类。在四个前沿数据集上的实验表明,PhishKey最高可达98.70% F1分数,对注入类对抗攻击具有强抵抗力,性能下降极小。

原文摘要 · Abstract (English)

Phishing attacks pose a significant cybersecurity threat, evolving rapidly to bypass detection mechanisms and exploit human vulnerabilities. This paper introduces PhishKey to address the challenges of adaptability, robustness, and efficiency. PhishKey is a novel phishing detection method using automatic feature extraction from hybrid sources. PhishKey combines character-level processing with Convolutional Neural Networks (CNN) for URL classification, and a Centroid-Based Key Component Phishing Extractor (CAPE) for HTML content at the word level. CAPE reduces noise and ensures complete sample processing avoiding crop operations on the input data. The predictions from both modules are integrated using a soft-voting ensemble to achieve more accurate and reliable classifications. Experimental evaluations on four state-of-the-art datasets demonstrate the effectiveness of PhishKey. It achieves up to 98.70% F1 Score and shows strong resistance to adversarial manipulations such as injection attacks with minimal performance degradation.

钓鱼检测CNN聚类网络安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。