arXiv:2602.09012cs.LGcs.AI2026-02被引 3

利用人机认知差异设计可扩展的新型验证码防御体系

Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense

  • 基于动态任务生成机制,构建可无限扩展的验证码实例
  • 在复杂逻辑题上实现90%通过率的攻击者仍无法破解
  • 适合对抗具备强推理能力的下一代智能代理

GUI智能体的快速发展已使传统验证码失效。尽管先前的OpenCaptchaWorld基准提供了多模态代理评估基础,但近期如Gemini3-Pro-High和GPT-5.2-Xhigh等强推理模型已能以高达90%的通过率破解复杂逻辑题(如'Bingo')。为此,我们提出下一代验证码(Next-Gen CAPTCHAs),一个可扩展的防御框架,旨在保护下一代网页安全。不同于静态数据集,本系统依托稳健的数据生成流水线,支持大规模、易扩展的评估,尤其对后端支持的类型,可生成近乎无限的验证码实例。通过利用人类与智能体在交互感知、记忆、决策与动作上的持续‘认知差距’,设计需适应性直觉而非精细规划的动态任务,重新建立生物用户与人工代理之间的可靠区分,为智能体时代提供可扩展且多样化的防御机制。

原文摘要 · Abstract (English)

The rapid evolution of GUI-enabled agents has rendered traditional CAPTCHAs obsolete. While previous benchmarks like OpenCaptchaWorld established a baseline for evaluating multimodal agents, recent advancements in reasoning-heavy models, such as Gemini3-Pro-High and GPT-5.2-Xhigh have effectively collapsed this security barrier, achieving pass rates as high as 90% on complex logic puzzles like "Bingo". In response, we introduce Next-Gen CAPTCHAs, a scalable defense framework designed to secure the next-generation web against the advanced agents. Unlike static datasets, our benchmark is built upon a robust data generation pipeline, allowing for large-scale and easily scalable evaluations, notably, for backend-supported types, our system is capable of generating effectively unbounded CAPTCHA instances. We exploit the persistent human-agent "Cognitive Gap" in interactive perception, memory, decision-making, and action. By engineering dynamic tasks that require adaptive intuition rather than granular planning, we re-establish a robust distinction between biological users and artificial agents, offering a scalable and diverse defense mechanism for the agentic era.

验证码人机对抗智能代理认知差距

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。