用强化学习+显式推理监督,让机器人自动破解验证码。
CaptchaMind: Training CAPTCHA Solvers via Reinforcement Learning with Explicit Reasoning Supervision

- 通过强化学习结合显式推理过程监督,提升解码能力。
- 在8类任务上平均成功率82.9%,真实场景达71.0%。
- 首个支持大规模训练的验证码基准,适合自动化研究者。
验证码广泛用于人机验证,常阻断智能代理在真实网络环境中的端到端自动化。破解现代验证码需强大的多步视觉推理与交互能力,但因缺乏大规模标注数据和过程级注释,基于训练的方法长期缺席。我们提出CaptchaBench,首个支持大规模训练的验证码基准,包含8个任务类别、16,000个程序生成样本,并提供区域级与过程级详细标注。系统评估显示,现有方法在需要细粒度视觉捕捉和区域对比的任务中表现持续失败。因此,我们提出CaptchaMind,一种基于强化学习并接受显式推理过程监督的求解器,在8个任务上实现82.9%的平均成功率,真实场景下达71.0%,显著优于所有无需闭源API的现有方法。
原文摘要 · Abstract (English)
CAPTCHAs are widely deployed as human verification mechanisms and frequently block intelligent agents from completing end-to-end automation in real-world web environments. Solving modern CAPTCHAs requires robust multi-step visual reasoning and interaction capabilities, yet training-based approaches have remained absent due to the lack of large-scale training data and process-level annotations. We introduce CaptchaBench, the first CAPTCHA benchmark designed to support large-scale training, comprising 16,000 programmatically generated samples across eight task categories with detailed region and process-level annotations. Systematic evaluation on CaptchaBench reveals that existing methods fail consistently on tasks requiring fine-grained visual detail capture and region-level comparison. We therefore present CaptchaMind, an RL-based solver trained with explicit reasoning process supervision, achieving 82.9% average success rate across eight tasks and 71.0% on real-world instances, substantially outperforming all existing methods without closed-source APIs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。