通过推理提升视觉语言模型破解验证码能力,准确率从21.9%升至83.9%
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
- 要求模型分步推理后再输出坐标,显著提升解题能力
- 新基准CAPTCHA-X涵盖7类验证码,含步骤与定位标注
- 提出5个推理评估指标,适用于高难度空间认知任务
CAPTCHA最初用于区分人类与机器人,现已成为评估视觉语言模型(VLMs)空间推理能力的真实世界基准。本文首次表明,分步推理对解决高难度空间推理任务(如CAPTCHA)至关重要,而当前主流商业VLMs(如Gemini、Claude、GPT等)仍难以有效应对,平均准确率仅约21.9%。研究发现,要求模型在生成最终坐标前进行分步推理,可显著提升解题准确率,凸显现有模型的严重不足。为此,我们提出CAPTCHA-X——首个包含推理过程的现实世界CAPTCHA基准,涵盖七类验证码(如Gobang、hCaptcha等),提供分步操作解法和定位标注。同时定义五项面向推理的评估指标,实现对模型推理能力的全面衡量。为进一步验证推理有效性,我们设计了一种基于代理的通用VLM框架,充分利用模型内在推理能力。该方法在五类高难度验证码上达到83.9%的平均准确率,显著超越现有基线,揭示了当前模型的局限性,并强调推理在推动未来视觉-空间挑战中的关键作用。
原文摘要 · Abstract (English)
CAPTCHA, originally designed to distinguish humans from robots, has evolved into a real-world benchmark for assessing the spatial reasoning capabilities of vision-language models. In this work, we first show that step-by-step reasoning is crucial for vision-language models (VLMs) to solve CAPTCHAs, which represent high-difficulty spatial reasoning tasks, and that current commercial vision-language models still struggle with such reasoning. In particular, we observe that most commercial VLMs (e.g., Gemini, Claude, GPT, etc.) fail to effectively solve CAPTCHAs and thus achieve low accuracy (around 21.9 percent). However, our findings indicate that requiring the model to perform step-by-step reasoning before generating the final coordinates can significantly enhance its solving accuracy, underscoring the severity of the gap. To systematically study this issue, we introduce CAPTCHA-X, the first real-world CAPTCHA benchmark with reasoning, covering seven categories of CAPTCHAs (such as Gobang, hCaptcha, etc.) with step-by-step action solutions and grounding annotations. We further define five reasoning-oriented metrics that enable a comprehensive evaluation of models reasoning capabilities. To validate the effectiveness of reasoning, we also propose a general agentic VLM-based framework that incorporates the models inherent reasoning abilities. Our method achieves state-of-the-art performance across five high-difficulty CAPTCHA types, with an average solving accuracy of 83.9 percent, substantially surpassing existing baselines. These results reveal the limitations of current models and highlight the importance of reasoning in advancing visual-spatial challenges in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。