首个跨模态验证码安全评测基准,揭示模型攻击下的漏洞规律
MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks
- 用统一视觉语言模型框架,适配多种验证码类型进行攻击测试
- 首次量化分析挑战复杂度、交互深度与模型破解率的关系
- 适合安全研究者和验证码设计者参考,推动防御体系升级
随着自动化攻击技术的快速发展,CAPTCHA 仍是抵御恶意机器人的重要防线。然而,现有验证码形式多样——从静态扭曲文本、混淆图像到交互式点击、滑动拼图和逻辑题——但学界缺乏统一、大规模、多模态的评测基准来严格评估其安全性。为此,我们提出 MCA-Bench,一个综合性且可复现的评测套件,将异构验证码类型整合到统一评估协议中。通过共享的视觉语言模型主干网络,我们为每类验证码微调专用破解代理,实现一致的跨模态评估。大量实验表明,MCA-Bench 能有效刻画现代验证码设计在不同攻击场景下的脆弱性谱系,并首次提供对挑战复杂度、交互深度与模型可解性之间关系的定量分析。基于这些发现,我们提出三条可操作的设计原则,识别关键开放问题,为系统化验证码加固、公平评测及社区协作奠定基础。数据集与代码已公开。
原文摘要 · Abstract (English)
As automated attack techniques rapidly advance, CAPTCHAs remain a critical defense mechanism against malicious bots. However, existing CAPTCHA schemes encompass a diverse range of modalities -- from static distorted text and obfuscated images to interactive clicks, sliding puzzles, and logic-based questions -- yet the community still lacks a unified, large-scale, multimodal benchmark to rigorously evaluate their security robustness. To address this gap, we introduce MCA-Bench, a comprehensive and reproducible benchmarking suite that integrates heterogeneous CAPTCHA types into a single evaluation protocol. Leveraging a shared vision-language model backbone, we fine-tune specialized cracking agents for each CAPTCHA category, enabling consistent, cross-modal assessments. Extensive experiments reveal that MCA-Bench effectively maps the vulnerability spectrum of modern CAPTCHA designs under varied attack settings, and crucially offers the first quantitative analysis of how challenge complexity, interaction depth, and model solvability interrelate. Based on these findings, we propose three actionable design principles and identify key open challenges, laying the groundwork for systematic CAPTCHA hardening, fair benchmarking, and broader community collaboration. Datasets and code are available online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。