arXiv:2604.04951cs.CRcs.AI2026-04

提出新型诈骗威胁模型,揭示AI如何通过制造信任操控人类决策。

Synthetic Trust Attacks: Modeling How Generative AI Manipulates Human Decisions in Social Engineering Fraud

  • 构建八阶段攻击框架,系统化描述从侦察到得手的全流程。
  • 实测显示人类对深度伪造识别率仅55.5%,远低于随机水平。
  • 建议从检测媒体转向干预决策,适合安全与反欺诈研究者。

想象收到一段视频通话,你的财务总监正与同事在会议中紧急要求你授权一笔机密转账,你照做后损失2500万美元。这并非虚构:2024年1月香港已发生类似事件,成为新一代诈骗的模板。本文提出‘合成信任攻击’(STAs)作为正式威胁类别,并引入STAM模型——一个覆盖从侦察到事后利用的八阶段操作框架。核心观点是:现有防御聚焦于合成媒体检测,但真实攻击面在于受害者的决策过程。当人类对深度伪造的识别准确率约为55.5%(仅略高于随机),而大语言模型诈骗代理达成46%的配合率(人类操作员仅18%),且完全绕过安全过滤时,感知层已失效。防御必须前移至决策层。本文提出五类信任线索分类法、17字段事件编码方案及可验证的四个假设,建立攻击结构与合规结果间的关联。同时将实践者开发的‘冷静、核查、确认’协议转化为可研究的决策层防御方法。真正的攻击面不是合成媒体,而是合成可信度。

原文摘要 · Abstract (English)

Imagine receiving a video call from your CFO, surrounded by colleagues, asking you to urgently authorise a confidential transfer. You comply. Every person on that call was fake, and you just lost $25 million. This is not a hypothetical. It happened in Hong Kong in January 2024, and it is becoming the template for a new generation of fraud. AI has not invented a new crime. It has industrialised an ancient one: the manufacture of trust. This paper proposes Synthetic Trust Attacks (STAs) as a formal threat category and introduces STAM, the Synthetic Trust Attack Model, an eight-stage operational framework covering the full attack chain from adversary reconnaissance through post-compliance leverage. The core argument is this: existing defenses target synthetic media detection, but the real attack surface is the victim's decision. When human deepfake detection accuracy sits at approximately 55.5%, barely above chance, and LLM scam agents achieve 46% compliance versus 18% for human operators while evading safety filters entirely, the perception layer has already failed. Defense must move to the decision layer. We present a five-category Trust-Cue Taxonomy, a reproducible 17-field Incident Coding Schema with a pilot-coded example, and four falsifiable hypotheses linking attack structure to compliance outcomes. The paper further operationalizes the author's practitioner-developed Calm, Check, Confirm protocol as a research-grade decision-layer defense. Synthetic credibility, not synthetic media, is the true attack surface of the AI fraud era.

AI诈骗信任操控决策防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。