arXiv:2411.02403cs.SIcs.AI2024-11被引 6

用心理学说服原理增强数据,提升诈骗短信识别能力

A Persuasion-Based Prompt Learning Approach to Improve Smishing Detection through Data Augmentation

  • 基于说服心理学设计提示词,生成更贴近真实诈骗短信的数据
  • 生成数据多样性与质量优于现有方法,显著提升模型检测精度
  • 特别适合训练大模型,对参数多的模型效果更明显

短信诈骗(smishing)因其对社会的负面影响而备受关注。尽管机器学习(ML)被广泛用于过滤此类信息,但其发展受限于标注数据稀缺——由于数据敏感性,公开可用的训练与评估数据极少。此外,诈骗短信与垃圾短信等社交工程攻击之间的细微相似性,进一步加剧了分类难度。为此,本文提出一种新颖的数据增强方法,采用少样本提示学习,并融入说服心理学原理。通过设计基于说服机制的提示词,所生成的数据能有效捕捉诈骗短信的关键特征,提升模型训练质量。在真实场景下的评估表明,该方法生成的数据更具多样性和高质量,显著增强了模型对诈骗短信微妙特征的识别能力。额外分析显示,该方法在参数量更大的模型上表现更优,证明其在训练大规模模型中的有效性。

原文摘要 · Abstract (English)

Smishing, which aims to illicitly obtain personal information from unsuspecting victims, holds significance due to its negative impacts on our society. In prior studies, as a tool to counteract smishing, machine learning (ML) has been widely adopted, which filters and blocks smishing messages before they reach potential victims. However, a number of challenges remain in ML-based smishing detection, with the scarcity of annotated datasets being one major hurdle. Specifically, given the sensitive nature of smishing-related data, there is a lack of publicly accessible data that can be used for training and evaluating ML models. Additionally, the nuanced similarities between smishing messages and other types of social engineering attacks such as spam messages exacerbate the challenge of smishing classification with limited resources. To tackle this challenge, we introduce a novel data augmentation method utilizing a few-shot prompt learning approach. What sets our approach apart from extant methods is the use of the principles of persuasion, a psychology theory which explains the underlying mechanisms of smishing. By designing prompts grounded in the persuasion principles, our augmented dataset could effectively capture various, important aspects of smishing messages, enabling ML models to be effectively trained. Our evaluation within a real-world context demonstrates that our augmentation approach produces more diverse and higher-quality smishing data instances compared to other cutting-edging approaches, leading to substantial improvements in the ability of ML models to detect the subtle characteristics of smishing messages. Moreover, our additional analyses reveal that the performance improvement provided by our approach is more pronounced when used with ML models that have a larger number of parameters, demonstrating its effectiveness in training large-scale ML models.

诈骗检测数据增强提示学习心理学应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。