测试大模型对抗欺诈诱导的多轮防御能力,发现角色扮演下漏洞明显。
Fraud-R1 : A Multi-Round Benchmark for Assessing the Robustness of LLM Against Augmented Fraud and Phishing Inducements
- 设计多轮评估流程,模拟真实欺诈场景中的信任建立与情绪操控。
- 15个大模型在假职位场景中表现差,角色扮演设置下错误率超60%。
- 中文模型比英文模型更易受骗,凸显多语言防护短板。
我们提出 Fraud-R1,一个用于评估大语言模型在动态真实场景中抵御网络欺诈与钓鱼攻击能力的基准。该基准包含8,564个来自钓鱼诈骗、虚假招聘、社交媒体和新闻的欺诈案例,分为5类主要欺诈类型。不同于以往基准,Fraud-R1引入多轮评估流程,考察模型在可信度建立、紧迫感营造和情绪操纵等不同阶段的抗骗能力。我们对15个大模型在两种设置下进行评估:1. 通用助手模式(Helpful-Assistant),提供一般决策辅助;2. 角色扮演模式(Role-play),模拟真实世界中基于代理的交互。结果表明,模型在防范欺诈诱导方面面临巨大挑战,尤其在角色扮演设置和虚假招聘场景中表现不佳。此外,我们观察到中英文模型间存在显著性能差距,凸显提升多语言欺诈检测能力的迫切需求。
原文摘要 · Abstract (English)
We introduce Fraud-R1, a benchmark designed to evaluate LLMs' ability to defend against internet fraud and phishing in dynamic, real-world scenarios. Fraud-R1 comprises 8,564 fraud cases sourced from phishing scams, fake job postings, social media, and news, categorized into 5 major fraud types. Unlike previous benchmarks, Fraud-R1 introduces a multi-round evaluation pipeline to assess LLMs' resistance to fraud at different stages, including credibility building, urgency creation, and emotional manipulation. Furthermore, we evaluate 15 LLMs under two settings: 1. Helpful-Assistant, where the LLM provides general decision-making assistance, and 2. Role-play, where the model assumes a specific persona, widely used in real-world agent-based interactions. Our evaluation reveals the significant challenges in defending against fraud and phishing inducement, especially in role-play settings and fake job postings. Additionally, we observe a substantial performance gap between Chinese and English, underscoring the need for improved multilingual fraud detection capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。