arXiv:2606.02449cs.AIcs.CL2026-06

测试智能体能否像人一样通过交互式验证码验证

HLL: Can Agents Cross Humanity's Last Line of Verification?

论文配图:HLL: Can Agents Cross Humanity's Last Line of Verification?
图 1 · 摘自论文原文
  • 设计真实界面压力下的交互式验证码评测基准
  • 现有智能体在复杂环境下验证成功率大幅下降
  • 适合评估智能体在受保护流程中替代人类的能力

多模态智能体日益被期望代表用户操作界面,引发核心部署问题:它们能否真正替代人类完成服务方刻意防范自动化的任务?验证码验证使这一问题具体化。它不仅是视觉谜题,更是账户创建、内容访问、表单提交等受保护操作前的人类验证屏障。我们提出「人类最后验证线(HLL)」,一个受控基准,利用交互式验证码评估智能体是否能通过基于现实的类人交互跨越此边界,而非仅依赖识别。HLL涵盖多种验证码交互,并引入可控的真实感压力因素,包括杂乱网页、更难任务变体及解题过程的有效动作追溯验证。我们在闭环图形界面环境中评估了八种前沿多模态智能体。结果表明,当前智能体在此人类替代边界上仍显脆弱:不同验证类型表现差异显著,真实界面条件下性能下降,且需有效动作轨迹支持时准确率进一步降低。该测试暴露了定位、动作校准、状态追踪和过程一致性方面的差距,为衡量多模态智能体在受保护现实工作流中接近人类替代的程度提供了具体基准。代码已公开于 https://github.com/XinhaoS0101/HLL。

原文摘要 · Abstract (English)

Multimodal agents are increasingly expected to operate interfaces on behalf of users, raising a central deployment question: can they truly substitute for humans in workflows that services deliberately protect against automation? CAPTCHA verification makes this question concrete. It is not merely a visual puzzle, but a human-verification boundary placed before account creation, content access, form submission, and other protected actions. We introduce \textbf{Humanity's Last Line of Verification (HLL)}, a controlled benchmark that uses interactive CAPTCHA verification to evaluate whether agents can cross this boundary through grounded, human-like interaction rather than recognition alone. HLL covers diverse CAPTCHA interactions and exposes agents to controlled realism stressors, including cluttered webpages, harder task variants, and trace-conditioned validation of the solving process. We evaluate eight frontier multimodal agents in a closed-loop GUI environment. The results show that current agents remain brittle at this human-substitution boundary: performance varies sharply across verification types, degrades under realistic interface conditions, and drops further when correct answers must be supported by valid action traces. By exposing gaps in localization, action calibration, state tracking, and process consistency, HLL provides a concrete testbed for measuring how close multimodal agents are to acting as human substitutes in protected real-world workflows. Our code is available at https://github.com/XinhaoS0101/HLL

智能体验证码人机替代评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。