arXiv:2509.08493cs.CRcs.AI2025-09被引 4

用大模型主动诱骗骗子,成功获取32%敏感信息

Send to which account? Evaluation of an LLM-based Scambaiting System

  • 用大语言模型自动对话诱骗真实骗子,收集威胁情报
  • 5个月间触达2600+骗子,32%成功获取账户等敏感信息
  • 生成回复70%符合人类操作员偏好,适合安全研究者参考

诈骗者正越来越多地利用生成式人工智能技术大规模制造逼真的钓鱼内容,加剧金融欺诈并破坏公众信任。传统防御手段如检测算法、用户培训和事后封禁虽重要,但难以瓦解诈骗者依赖的洗钱账户和加密货币钱包等基础设施。为此,一种新兴的主动策略是使用对话式诱饵系统与骗子互动以获取可操作的威胁情报。本文首次对基于大语言模型(LLM)的诱骗系统进行了大规模真实世界评估。在为期五个月的部署中,该系统与2600多个真实骗子展开了超过2600次互动,生成了超过18,700条消息数据集。系统实现了约32%的信息披露率(IDR),成功提取了包括收款账户在内的敏感金融信息。同时,系统保持约70%的人类接受率(HAR),表明其生成回复与人工操作员偏好高度一致。分析还揭示关键挑战:初始接触成功率仅48.7%,多数骗子未回应首个诱导消息。这些发现凸显了进一步优化的必要性,并为自动化诱骗系统的设计提供了可操作洞察。

原文摘要 · Abstract (English)

Scammers are increasingly harnessing generative AI(GenAI) technologies to produce convincing phishing content at scale, amplifying financial fraud and undermining public trust. While conventional defenses, such as detection algorithms, user training, and reactive takedown efforts remain important, they often fall short in dismantling the infrastructure scammers depend on, including mule bank accounts and cryptocurrency wallets. To bridge this gap, a proactive and emerging strategy involves using conversational honeypots to engage scammers and extract actionable threat intelligence. This paper presents the first large-scale, real-world evaluation of a scambaiting system powered by large language models (LLMs). Over a five-month deployment, the system initiated over 2,600 engagements with actual scammers, resulting in a dataset of more than 18,700 messages. It achieved an Information Disclosure Rate (IDR) of approximately 32%, successfully extracting sensitive financial information such as mule accounts. Additionally, the system maintained a Human Acceptance Rate (HAR) of around 70%, indicating strong alignment between LLM-generated responses and human operator preferences. Alongside these successes, our analysis reveals key operational challenges. In particular, the system struggled with engagement takeoff: only 48.7% of scammers responded to the initial seed message sent by defenders. These findings highlight the need for further refinement and provide actionable insights for advancing the design of automated scambaiting systems.

AI反诈大模型应用安全研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。