arXiv:2605.17245cs.NIcs.LG2026-05

用机器学习精准识别电信诈骗,准确率超99.9%。

An Efficient Machine Learning-based Framework for Detection and Prevention of Frauds in Telecom Networks

  • 基于通话记录数据,用随机森林与XGBoost模型进行欺诈检测。
  • 随机森林准确率达99.9%,误判极少,表现最优。
  • 适合电信安全团队快速部署,提升反诈能力。

电信诈骗是全球范围内导致重大经济损失并威胁通信系统可靠性的严重问题。本文基于呼叫详单(CDR)数据集评估了人工智能驱动模型在电信网络欺诈检测中的性能。研究使用包含101,174条客户记录、17个属性的Telecom CDR数据集,其中包含8,830例欺诈案例。特征预处理包括缺失值处理、最小-最大缩放及通过SMOTE技术实现数据平衡。采用随机森林(RF)和XGBoost模型进行预测分析,以F1-score、ROC AUC、召回率、准确率、时间与精确率作为评估指标。结果显示,随机森林准确率达到99.9%,精确率、召回率与F1-score均为99.9%;XGBoost为99.7%。对比包括DBSCAN、RoBERTa、K-means、GNN与BERT在内的多种模型,随机森林表现最佳。研究证实所提框架能高效、准确地检测欺诈行为,具备实际应用价值。

原文摘要 · Abstract (English)

Telecommunication fraud is an acute problem that leads to substantial material losses and compromises the reliability of telecom systems worldwide. Only effective and efficient detection mechanisms can help to deal with these threats, though there are certain shifts in the approaches to fraud detection. This paper evaluates the performance of AI-driven models for fraud detection in telecommunication networks using Call Detail Record (CDR) datasets. This study focuses on fraud detection in telecom networks using the Telecom CDR dataset, which contains 101,174 customer records with 17 attributes, including 8,830 fraud cases. In feature preprocessing, missing values were dealt with, followed by data scaling using Min-Max scaling and data balancing using the SMOTE technique. The dataset was trained for predictive analysis using Random Forest (RF) and XGBoost models. F1-score, ROC AUC, recall, accuracy, time, and precision were used as indicators with which to compare performance of the two models. RF recorded a high level of accuracy at 99.9% while XGBoost at 99.7%. Findings show that the suggested framework successfully detects fraud with few misclassifications. Several machine learning models were evaluated and contrasted, such as RF, XGBoost, DBSCAN, RoBERTa, and K-means. Among all the models, RF was seen to give the highest performance with an accuracy of 99.9% and precision of 99.9%, recall of 99.9% and F1-score of 99.9%, XGBoost, GNN and BERT. The findings emphasize RF as the most effective model for detecting fraudulent activities in telecom networks, ensuring robust and reliable prevention of fraud.

电信诈骗机器学习欺诈检测随机森林

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。