首个专为大视觉语言模型设计的验证码评测基准,全面评估其解码能力。
CAPTURE: A Benchmark and Evaluation for LVLMs in CAPTCHA Resolving
- 构建覆盖4类25子类的多源验证码数据集,专为LVLM优化标签体系。
- 实测显示当前主流LVLM在真实场景验证码上表现不佳,准确率普遍低于60%。
- 适合研究多模态模型鲁棒性、安全验证机制及对抗样本防御的学者使用。
得益于强大的多模态对齐策略,大视觉语言模型(LVLM)能够模拟人类的视觉与推理能力,如破解验证码。然而,现有的基于视觉验证码的评测基准仍存在局限:以往研究多根据自身目标定制数据集,导致覆盖不全,且缺乏针对LVLM的专用评测基准。为此,本文首次提出CAPTURE——一个专为LVLM设计的验证码评测基准,全称为CAPTCHA for Testing Under Real-world Experiments。该基准涵盖31家厂商提供的4大类、25个子类验证码,具有高度多样性,支持对LVLM性能进行多维度、系统性评估。CAPTURE具备丰富的类别覆盖、大规模数据量及专为LVLM设计的标注体系,在数据全面性和标签相关性方面填补了现有研究空白。在该基准上的测试表明,当前主流LVLM在真实场景验证码破解任务中表现较差,平均准确率不足60%。
原文摘要 · Abstract (English)
Benefiting from strong and efficient multi-modal alignment strategies, Large Visual Language Models (LVLMs) are able to simulate human visual and reasoning capabilities, such as solving CAPTCHAs. However, existing benchmarks based on visual CAPTCHAs still face limitations. Previous studies, when designing benchmarks and datasets, customized them according to their research objectives. Consequently, these benchmarks cannot comprehensively cover all CAPTCHA types. Notably, there is a dearth of dedicated benchmarks for LVLMs. To address this problem, we introduce a novel CAPTCHA benchmark for the first time, named CAPTURE CAPTCHA for Testing Under Real-world Experiments, specifically for LVLMs. Our benchmark encompasses 4 main CAPTCHA types and 25 sub-types from 31 vendors. The diversity enables a multi-dimensional and thorough evaluation of LVLM performance. CAPTURE features extensive class variety, large-scale data, and unique LVLM-tailored labels, filling the gaps in previous research in terms of data comprehensiveness and labeling pertinence. When evaluated by this benchmark, current LVLMs demonstrate poor performance in solving CAPTCHAs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。