用文本标准化提升短信诈骗识别率,准确率达96.2%
Machine Learning Driven Smishing Detection Framework for Mobile Security
- 通过文本归一化将非标准短信转为规范表达
- 在公开数据集上实现96.2%检测准确率
- 适合移动安全防护与反诈骗系统开发
智能手机在通信、金融交易和个人数据管理中的广泛应用使其成为网络攻击的主要目标,尤其是通过短信进行的诈骗(smishing)。传统检测方法难以应对短信语言中常见的缩写、俚语和简写形式。本文提出一种增强的内容驱动式短信诈骗检测框架,利用先进的文本归一化技术提升机器学习分类器的性能,特别是朴素贝叶斯分类器。该方法将非标准文本转换为标准形式,显著提升了区分诈骗短信与正常短信的能力。实验结果基于公开数据集验证,检测准确率达96.2%,误报率为3.87%,漏报率为2.85%,显著优于现有方法,为移动环境下的复杂诈骗威胁提供了可靠解决方案。
原文摘要 · Abstract (English)
The increasing reliance on smartphones for communication, financial transactions, and personal data management has made them prime targets for cyberattacks, particularly smishing, a sophisticated variant of phishing conducted via SMS. Despite the growing threat, traditional detection methods often struggle with the informal and evolving nature of SMS language, which includes abbreviations, slang, and short forms. This paper presents an enhanced content-based smishing detection framework that leverages advanced text normalization techniques to improve detection accuracy. By converting nonstandard text into its standardized form, the proposed model enhances the efficacy of machine learning classifiers, particularly the Naive Bayesian classifier, in distinguishing smishing messages from legitimate ones. Our experimental results, validated on a publicly available dataset, demonstrate a detection accuracy of 96.2%, with a low False Positive Rate of 3.87% and False Negative Rate of 2.85%. This approach significantly outperforms existing methodologies, providing a robust solution to the increasingly sophisticated threat of smishing in the mobile environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。