arXiv:2506.10685cs.CVcs.CR2025-06被引 19

用语义引导生成难以被机器识别的验证码,提升安全性和通用性。

Defensive Adversarial CAPTCHA: A Semantics-Driven Framework for Natural Adversarial Example Generation

  • 基于大语言模型和扩散机制,按语义生成高保真对抗样本。
  • 在白盒攻击中实现精准语义对齐,黑盒攻击下准确率超90%。
  • 适合需要对抗未知模型攻击的场景,如在线安全验证系统。

传统验证码易受深度神经网络攻击。现有方法依赖原始图像特征,导致失真严重且无法在无初始图像时使用。为此,我们提出无源对抗验证码(DAC),通过攻击者指定的语义信息生成高质量对抗样本。利用大语言模型增强验证码多样性和语义丰富性。针对白盒目标攻击,引入两个交替引导的潜在噪声变量,在扩散步骤中实现鲁棒逆向;梯度引导与潜在变量优化的协同确保生成样本在语义一致性与攻击有效性上表现最优。针对黑盒无目标攻击,提出双路径无源对抗验证码(BP-DAC),采用多模态梯度与双路径优化策略,实现高效误分类。实验表明,BP-DAC生成的防御型对抗验证码能抵御多数未知模型攻击,且对人类与深度神经网络均不可区分。

原文摘要 · Abstract (English)

Traditional CAPTCHA (Completely Automated Public Turing Test to Tell Computers and Humans Apart) schemes are increasingly vulnerable to automated attacks powered by deep neural networks (DNNs). Existing adversarial attack methods often rely on the original image characteristics, resulting in distortions that hinder human interpretation and limit their applicability in scenarios where no initial input images are available. To address these challenges, we propose the Unsourced Adversarial CAPTCHA (DAC), a novel framework that generates high-fidelity adversarial examples guided by attacker-specified semantics information. Leveraging a Large Language Model (LLM), DAC enhances CAPTCHA diversity and enriches the semantic information. To address various application scenarios, we examine the white-box targeted attack scenario and the black box untargeted attack scenario. For target attacks, we introduce two latent noise variables that are alternately guided in the diffusion step to achieve robust inversion. The synergy between gradient guidance and latent variable optimization achieved in this way ensures that the generated adversarial examples not only accurately align with the target conditions but also achieve optimal performance in terms of distributional consistency and attack effectiveness. In untargeted attacks, especially for black-box scenarios, we introduce bi-path unsourced adversarial CAPTCHA (BP-DAC), a two-step optimization strategy employing multimodal gradients and bi-path optimization for efficient misclassification. Experiments show that the defensive adversarial CAPTCHA generated by BP-DAC is able to defend against most of the unknown models, and the generated CAPTCHA is indistinguishable to both humans and DNNs.

验证码对抗样本扩散模型LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。