无需查询即可攻击视频模型,通过修改特征图实现隐蔽对抗攻击
FeatureFool: Zero-Query Fooling of Video Models via Feature Map
- 利用深度网络提取的特征图直接扰动视频特征空间
- 零查询攻击成功率超70%,且可绕过视频大模型检测
- 生成视频质量高,感知上几乎不可察觉,适合真实场景攻击
深度神经网络的脆弱性已得到初步验证。现有黑盒对抗攻击通常需要与模型多次交互,消耗大量查询,不适用于现实场景,也难以扩展到新兴的视频大模型(Video-LLMs)。此外,视频领域尚无攻击直接利用特征图来转移干净视频的特征空间。为此,我们提出 FeatureFool——一种隐蔽、零查询、面向视频领域的黑盒攻击方法,通过提取深度网络信息来改变干净视频的特征空间。与依赖迭代查询的方法不同,FeatureFool直接利用DNN提取的信息完成零查询攻击,该方法在视频领域前所未见。实验表明,FeatureFool在不进行任何查询的情况下,对传统视频分类器的攻击成功率超过70%。得益于特征图的迁移性,该方法还能生成有害内容以绕过视频大模型识别。此外,生成的对抗视频在SSIM、PSNR和时间一致性方面表现良好,攻击几乎无法被感知。本文可能包含暴力或敏感内容。
原文摘要 · Abstract (English)
The vulnerability of deep neural networks (DNNs) has been preliminarily verified. Existing black-box adversarial attacks usually require multi-round interaction with the model and consume numerous queries, which is impractical in the real-world and hard to scale to recently emerged Video-LLMs. Moreover, no attack in the video domain directly leverages feature maps to shift the clean-video feature space. We therefore propose FeatureFool, a stealthy, video-domain, zero-query black-box attack that utilizes information extracted from a DNN to alter the feature space of clean videos. Unlike query-based methods that rely on iterative interaction, FeatureFool performs a zero-query attack by directly exploiting DNN-extracted information. This efficient approach is unprecedented in the video domain. Experiments show that FeatureFool achieves an attack success rate above 70\% against traditional video classifiers without any queries. Benefiting from the transferability of the feature map, it can also craft harmful content and bypass Video-LLM recognition. Additionally, adversarial videos generated by FeatureFool exhibit high quality in terms of SSIM, PSNR, and Temporal-Inconsistency, making the attack barely perceptible. This paper may contain violent or explicit content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。