用空间推理设计新验证码,人轻松AI难破。
Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine Differentiation
- 生成需几何推理的动态问题,利用人机差异防破解
- 10个顶尖模型最高仅31.0%正确率,人类显著领先
- 适合安全验证与评估AI空间认知能力的场景
在线服务依赖验证码作为防范自动化滥用的第一道防线,但多模态大模型的发展已削弱传统以文本识别或2D图像理解为主的验证码有效性。为此,我们提出Spatial CAPTCHA,一种基于人与多模态大模型在空间推理能力上根本差异的新型人机验证框架。该系统通过程序化生成机制,设计需几何推理、视角转换、遮挡处理和心理旋转的动态问题,这些能力对人类直观易懂,却对当前最先进的(SOTA)AI系统极具挑战。系统采用约束式难度控制、自动正确性验证及人机协同验证,确保可扩展性、鲁棒性与适应性。在对应基准Spatial-CAPTCHA-Bench上的评估显示,人类远超10个SOTA MLLMs,最佳模型仅达31.0%的Pass@1准确率。进一步对比Google reCAPTCHA,验证了其在安全防护与人工智能空间推理诊断方面的双重价值。
原文摘要 · Abstract (English)
Online services rely on CAPTCHAs as a first line of defense against automated abuse, yet recent advances in multi-modal large language models (MLLMs) have eroded the effectiveness of conventional designs that focus on text recognition or 2D image understanding. To address this challenge, we present Spatial CAPTCHA, a novel human-verification framework that leverages fundamental differences in spatial reasoning between humans and MLLMs. Unlike existing CAPTCHAs which rely on low-level perception tasks that are vulnerable to modern AI, Spatial CAPTCHA generates dynamic questions requiring geometric reasoning, perspective-taking, occlusion handling, and mental rotation. These skills are intuitive for humans but difficult for state-of-the-art (SOTA) AI systems. The system employs a procedural generation pipeline with constraint-based difficulty control, automated correctness verification, and human-in-the-loop validation to ensure scalability, robustness, and adaptability. Evaluation on a corresponding benchmark, Spatial-CAPTCHA-Bench, demonstrates that humans vastly outperform 10 state-of-the-art MLLMs, with the best model achieving only 31.0% Pass@1 accuracy. Furthermore, we compare Spatial CAPTCHA with Google reCAPTCHA, which confirms its effectiveness as both a security mechanism and a diagnostic tool for spatial reasoning in AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。