揭露大模型在诈骗检测中的漏洞,提出增强鲁棒性的方法。
Exposing LLM Vulnerabilities: Adversarial Scam Detection and Performance
- 构建包含原始与对抗性诈骗消息的细粒度数据集
- 发现对抗样本使大模型误判率显著升高
- 提出策略提升大模型在诈骗检测中的抗攻击能力
大型语言模型(LLMs)能否准确预测诈骗信息?本文研究了大模型在面对对抗性诈骗消息时的脆弱性。通过构建一个包含细粒度标签的综合性数据集,涵盖原始及对抗性诈骗消息,将传统的二分类诈骗检测任务扩展为更细致的诈骗类型识别。分析表明,对抗样本利用了大模型的内在弱点,导致高误判率。我们评估了大模型在这些对抗性消息上的表现,并提出了提升其鲁棒性的策略。
原文摘要 · Abstract (English)
Can we trust Large Language Models (LLMs) to accurately predict scam? This paper investigates the vulnerabilities of LLMs when facing adversarial scam messages for the task of scam detection. We addressed this issue by creating a comprehensive dataset with fine-grained labels of scam messages, including both original and adversarial scam messages. The dataset extended traditional binary classes for the scam detection task into more nuanced scam types. Our analysis showed how adversarial examples took advantage of vulnerabilities of a LLM, leading to high misclassification rate. We evaluated the performance of LLMs on these adversarial scam messages and proposed strategies to improve their robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。