AI agents可生成逼真诈骗电话,突破现有安全防护
ScamAgents: How AI Agents Can Simulate Human-Level Scam Calls
- 构建多轮对话的自主代理,动态应对用户回应
- 现有安全机制对这类代理基本无效,可绕过过滤
- 可将诈骗脚本转为真人语音,实现自动化诈骗
大型语言模型(LLMs)展现出出色的流畅性和推理能力,但其被滥用的潜在风险引发广泛关注。本文提出ScamAgent,一种基于LLM的自主多轮对话代理,能够生成高度逼真的诈骗电话脚本,模拟真实世界诈骗场景。与以往关注单次提示滥用的研究不同,ScamAgent具备对话记忆,能根据模拟用户回复动态调整策略,并在多轮对话中采用欺骗性说服技巧。我们发现,当前的LLM安全防护措施,包括拒绝机制和内容过滤,对这类代理威胁几乎无效。即使模型在提示层面有强防护,仍可通过分解、伪装或分步输入的方式在代理框架内绕过。我们进一步展示了如何利用现代文本转语音系统将诈骗脚本转化为逼真语音,完成全自动诈骗流程。研究揭示了亟需多轮安全审计、代理级控制框架以及针对生成式AI驱动对话欺骗的新检测与阻断方法。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated impressive fluency and reasoning capabilities, but their potential for misuse has raised growing concern. In this paper, we present ScamAgent, an autonomous multi-turn agent built on top of LLMs, capable of generating highly realistic scam call scripts that simulate real-world fraud scenarios. Unlike prior work focused on single-shot prompt misuse, ScamAgent maintains dialogue memory, adapts dynamically to simulated user responses, and employs deceptive persuasion strategies across conversational turns. We show that current LLM safety guardrails, including refusal mechanisms and content filters, are ineffective against such agent-based threats. Even models with strong prompt-level safeguards can be bypassed when prompts are decomposed, disguised, or delivered incrementally within an agent framework. We further demonstrate the transformation of scam scripts into lifelike voice calls using modern text-to-speech systems, completing a fully automated scam pipeline. Our findings highlight an urgent need for multi-turn safety auditing, agent-level control frameworks, and new methods to detect and disrupt conversational deception powered by generative AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。