arXiv:2606.21037cs.CRcs.CL2026-06被引 1

测试21个大模型发现,它们比人类更易被骗,且明知是陷阱也常去碰。

Honeyquest for LLMs: Rethinking Cyber Deception for AI Attackers

  • 用自动化框架对比大模型与人类对网络诱骗的反应
  • 所有大模型被骗率均高于人类,73.4%明知陷阱仍攻击
  • 现有针对人类的防御策略不适用于大模型,需新防御设计

基于人类中心假设的网络欺骗实证基础,面临自主型AI攻击者的挑战。为此,我们引入适配自Honeyquest的自动化评估框架,大规模评估LLM攻击者判断力。研究涵盖21个大模型,来自10家厂商,覆盖8B至超1T参数规模,包含开源与闭源模型。在相同174个侦察查询下,评估其表现(共10,962条响应),并与47名人类基线对比。结果揭示三个关键发现:(1)所有模型被骗率显著高于人类;(2)人类中观察到的注意力分散防御效应在大模型中统计上不存在;(3)存在关键的认知-行动鸿沟:73.4%的大模型在推理中识别出陷阱却仍攻击,其中48.5%明确识别后仍攻击,24.8%误判后攻击。大模型推理中的陷阱识别无法预测其实际行为(Spearman r = +0.08, p = 0.73)。结论表明,人类中心的欺骗假设不适用于AI攻击者,亟需构建面向AI的主动防御体系。

原文摘要 · Abstract (English)

The empirical foundation of cyber deception relies on human-centered hypotheses, but the rapid emergence of autonomous, AI-enabled attackers challenges whether this foundation transfers to AI agents. To address this, we introduce an automated evaluation framework adapted from the Honeyquest instrument to assess LLM attacker judgment at scale. Our 21-LLM cohort spanned 10 providers, diverse architectures and specializations, open- and closed-weight models, and parameter scales from 8B to over 1T. We evaluated the performance of this LLM cohort (yielding 10,962 responses) against the 47-participant human baseline across an identical set of 174 reconnaissance queries. Our empirical evaluation reveals three key findings that establish LLMs as a distinct attacker class: (1) every model in our cohort falls for deceptive traps at a significantly higher rate than human attackers; (2) the defensive attention-diversion effect observed in humans is statistically absent in our LLM cohort; and (3) a critical recognition-action gap, where LLMs successfully articulate trap recognition in their reasoning but exploit the deceptive elements anyway 73.4% of the time; 48.5% of aware-on-deceptive responses correctly identify the trap and exploit it anyway, while 24.8% exploit after misidentifying the deceptive line. Across the 21 models, trap recognition in reasoning text did not predict fell-for-trap behavior (Spearman $r = +0.08$, $p = 0.73$). Ultimately, these findings demonstrate that human-centered deception hypotheses do not reliably transfer to AI attackers, highlighting the critical need for new research into AI-native active defense frameworks.

AI安全网络欺骗大模型评估主动防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。