用关键词特征提升钓鱼网址检测准确率,小数据集效果更显著。
Enhance the machine learning algorithm performance in phishing detection with keyword features
- 从网址中提取关键词特征,与传统特征结合提升模型性能。
- 平均降低30%分类错误率,小数据集提升更明显。
- 不依赖第三方服务,最佳结果达99.68%准确率,适合安全检测场景。
近年来互联网钓鱼攻击显著增加。攻击者创建外观类似正规网站的恶意网站,以窃取用户信息,导致敏感数据泄露和财务损失。早期识别此类网址的URL至关重要。已有研究提出多种机器学习算法区分钓鱼与合法网址。本文从特征选择角度提升算法性能,提出一种将关键词特征融入传统特征的新方法。该方法应用于多个传统机器学习算法,实验表明其有效且实用:在大规模数据集上平均降低30%分类错误率;小数据集上提升更显著。该方法仅从URL中提取信息,无需依赖第三方服务。采用该方法的最佳模型达到99.68%的准确率。
原文摘要 · Abstract (English)
Recently, we can observe a significant increase of the phishing attacks in the Internet. In a typical phishing attack, the attacker sets up a malicious website that looks similar to the legitimate website in order to obtain the end-users' information. This may cause the leakage of the sensitive information and the financial loss for the end-users. To avoid such attacks, the early detection of these websites' URLs is vital and necessary. Previous researchers have proposed many machine learning algorithms to distinguish the phishing URLs from the legitimate ones. In this paper, we would like to enhance these machine learning algorithms from the perspective of feature selection. We propose a novel method to incorporate the keyword features with the traditional features. This method is applied on multiple traditional machine learning algorithms and the experimental results have shown this method is useful and effective. On average, this method can reduce the classification error by 30% for the large dataset. Moreover, its enhancement is more significant for the small dataset. In addition, this method extracts the information from the URL and does not rely on the additional information provided by the third-part service. The best result for the machine learning algorithm using our proposed method has achieved the accuracy of 99.68%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。