构建新数据集与方法,提升恶意网址分类准确率与实时性
A New Dataset and Methodology for Malicious URL Classification
- 构建多类恶意网址数据集DeepURLBench,支持良性、钓鱼、恶意三类区分
- 引入DNS特征增强模型,性能显著提升且保持实时推理速度
- 改进URLNet架构,为网络安全场景提供高效高精度分类方案
恶意网址分类是网络安全的关键环节,可防御网络威胁。尽管深度学习在该领域前景广阔,但其发展受限于两大问题:缺乏全面的开源数据集,以及现有模型难以兼顾实时性与高性能。为此,我们提出一个新型多类恶意网址分类数据集DeepURLBench,涵盖良性、钓鱼和恶意三类,经严格清洗与结构化处理,优于现有数据集。多类分类策略显著提升深度学习模型表现,相较传统二分类更优。同时,我们对基于字符串的网址分类器进行改进,应用于URLNet,关键在于融合DNS衍生特征,显著增强模型能力,在保持实时运行效率的前提下实现性能突破,适用于实际网络安全部署。
原文摘要 · Abstract (English)
Malicious URL (Uniform Resource Locator) classification is a pivotal aspect of Cybersecurity, offering defense against web-based threats. Despite deep learning's promise in this area, its advancement is hindered by two main challenges: the scarcity of comprehensive, open-source datasets and the limitations of existing models, which either lack real-time capabilities or exhibit suboptimal performance. In order to address these gaps, we introduce a novel, multi-class dataset for malicious URL classification, distinguishing between benign, phishing and malicious URLs, named DeepURLBench. The data has been rigorously cleansed and structured, providing a superior alternative to existing datasets. Notably, the multi-class approach enhances the performance of deep learning models, as compared to a standard binary classification approach. Additionally, we propose improvements to string-based URL classifiers, applying these enhancements to URLNet. Key among these is the integration of DNS-derived features, which enrich the model's capabilities and lead to notable performance gains while preserving real-time runtime efficiency-achieving an effective balance for cybersecurity applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。