系统梳理AI与NLP在在线诈骗检测中的应用现状
Application of AI-based Models for Online Fraud Detection and Analysis
- 采用PRISMA-ScR框架筛选223篇论文,覆盖16类诈骗行为
- 发现现有模型针对特定诈骗类型训练,泛化能力差
- 指出数据局限与评估偏差问题,对政策制定有参考价值
诈骗不仅造成经济损失,还给受害者带来心理和身体伤害。随着在线通信技术的发展,在线诈骗在庞大网络中迅速蔓延,欺诈者越来越多地利用这些渠道进行欺骗。随着AI等技术的进步,诈骗规模可能进一步扩大,例如使用ChatGPT等语言生成模型制造深度伪造的钓鱼攻击。然而,关于人工智能在检测和分析在线诈骗中的应用研究仍不充分。本文开展了一项系统文献综述(SLR),聚焦于AI与自然语言处理(NLP)技术在在线诈骗检测中的应用。遵循PRISMA-ScR协议,纳入标准包括与在线诈骗相关、使用文本数据及采用AI方法。共筛查2,457篇学术记录,350篇符合标准,最终纳入223篇。报告了当前最先进的NLP技术在各类在线诈骗分析中的应用情况;训练数据来源;构建的NLP算法与模型;以及用于模型评估的性能指标。研究发现,当前研究分散于多种诈骗活动,识别出16种研究人员关注的诈骗类型。该综述提升了学术界对基于AI的在线诈骗检测方法的理解,并为政策制定者、执法机构和企业提供防范建议。结论指出,针对特定诈骗类型的专注缺乏泛化性,需为不同诈骗类型开发多个模型。诈骗形式不断演变,导致基于过时数据训练的模型效果下降。同时,研究还揭示了数据限制、训练偏差报告不足以及评估指标选择性呈现等问题,可能导致模型评估出现潜在偏差。
原文摘要 · Abstract (English)
Fraud is a prevalent offence that extends beyond financial loss, causing psychological and physical harm to victims. The advancements in online communication technologies alowed for online fraud to thrive in this vast network, with fraudsters increasingly using these channels for deception. With the progression of technologies like AI, there is a growing concern that fraud will scale up, using sophisticated methods, like deep-fakes in phishing campaigns, all generated by language generation models like ChatGPT. However, the application of AI in detecting and analyzing online fraud remains understudied. We conduct a Systematic Literature Review on AI and NLP techniques for online fraud detection. The review adhered the PRISMA-ScR protocol, with eligibility criteria including relevance to online fraud, use of text data, and AI methodologies. We screened 2,457 academic records, 350 met our eligibility criteria, and included 223. We report the state-of-the-art NLP techniques for analysing various online fraud categories; the training data sources; the NLP algorithms and models built; and the performance metrics employed for model evaluation. We find that current research on online fraud is divided into various scam activitiesand identify 16 different frauds that researchers focus on. This SLR enhances the academic understanding of AI-based detection methods for online fraud and offers insights for policymakers, law enforcement, and businesses on safeguarding against such activities. We conclude that focusing on specific scams lacks generalization, as multiple models are required for different fraud types. The evolving nature of scams limits the effectiveness of models trained on outdated data. We also identify issues in data limitations, training bias reporting, and selective presentation of metrics in model performance reporting, which can lead to potential biases in model evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。