首个测试多模态大模型智能体解验证码能力的在线平台
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
- 构建20类共225个动态验证码的在线评测平台
- 人类解题准确率达93.3%,顶尖模型最高仅40.0%
- 引入认知与操作步骤数的新评估指标,适合研究交互推理
CAPTCHA是部署网络智能体于真实应用中的关键瓶颈,常阻碍其完成端到端自动化任务。尽管现代多模态大模型在静态感知任务中表现优异,但对其处理交互式、多步推理挑战(如CAPTCHA)的能力仍缺乏系统评估。为此,我们提出Open CaptchaWorld,首个专为评估多模态大模型智能体视觉推理与交互能力而设计的网页基准平台,涵盖20类现代CAPTCHA,共计225个实例,并引入新指标——CAPTCHA Reasoning Depth,量化每个谜题所需的认知与操作步骤数。实验表明,人类表现接近完美(93.3%准确率),而最先进多模态大模型智能体表现显著不足,最高成功率为40.0%(Browser-Use Openai-o3),远低于人类水平。该平台可有效诊断当前多模态智能体的局限性,推动更鲁棒的多模态推理系统发展。代码与数据已公开。
原文摘要 · Abstract (English)
CAPTCHAs have been a critical bottleneck for deploying web agents in real-world applications, often blocking them from completing end-to-end automation tasks. While modern multimodal LLM agents have demonstrated impressive performance in static perception tasks, their ability to handle interactive, multi-step reasoning challenges like CAPTCHAs is largely untested. To address this gap, we introduce Open CaptchaWorld, the first web-based benchmark and platform specifically designed to evaluate the visual reasoning and interaction capabilities of MLLM-powered agents through diverse and dynamic CAPTCHA puzzles. Our benchmark spans 20 modern CAPTCHA types, totaling 225 CAPTCHAs, annotated with a new metric we propose: CAPTCHA Reasoning Depth, which quantifies the number of cognitive and motor steps required to solve each puzzle. Experimental results show that humans consistently achieve near-perfect scores, state-of-the-art MLLM agents struggle significantly, with success rates at most 40.0% by Browser-Use Openai-o3, far below human-level performance, 93.3%. This highlights Open CaptchaWorld as a vital benchmark for diagnosing the limits of current multimodal agents and guiding the development of more robust multimodal reasoning systems. Code and Data are available at this https URL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。