arXiv:2505.18233cs.LGcs.AI2025-05被引 7

通过多信号融合提升诈骗短信识别率,准确率达97.89%

POSTER: A Multi-Signal Model for Detecting Evasive Smishing

  • 融合语义、结构、字符风格和上下文嵌入多信号检测
  • 在超8万条数据上实现97.89%准确率,AUC达99.73%
  • 适合安全团队部署,尤其适用于区域化诈骗防范

短信钓鱼(Smishing)通过文化适配、简洁且具有欺骗性的信息,模仿合法通信,对移动用户构成日益严重的威胁,可能导致敏感数据或财务损失。为此,我们提出一种多通道短信钓鱼检测模型,结合国家特定语义标签、结构模式标签、字符级风格特征及上下文短语嵌入。我们整理并重新标注了来自五个数据集的超过84,000条消息,其中包含24,086条钓鱼样本。统一架构在测试中达到97.89%准确率、0.963的F1分数和99.73%的AUC,显著优于单流模型,充分展现了多信号学习在鲁棒且区域感知的钓鱼检测中的有效性。

原文摘要 · Abstract (English)

Smishing, or SMS-based phishing, poses an increasing threat to mobile users by mimicking legitimate communications through culturally adapted, concise, and deceptive messages, which can result in the loss of sensitive data or financial resources. In such, we present a multi-channel smishing detection model that combines country-specific semantic tagging, structural pattern tagging, character-level stylistic cues, and contextual phrase embeddings. We curated and relabeled over 84,000 messages across five datasets, including 24,086 smishing samples. Our unified architecture achieves 97.89% accuracy, an F1 score of 0.963, and an AUC of 99.73%, outperforming single-stream models by capturing diverse linguistic and structural cues. This work demonstrates the effectiveness of multi-signal learning in robust and region-aware phishing.

短信诈骗多信号检测安全防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。