用图像攻击视频质量评估,跨模态攻击成功率高。
Cross-Modal Transferable Image-to-Video Attack on Video Quality Metrics
- 从图像模型生成扰动,迁移到视频质量评估模型
- 在三个黑盒视频评估模型上成功攻击率超90%
- 结合CLIP提升低层语义迁移能力,适合安全评测研究
近期研究表明,现代图像与视频质量评估(IQA/VQA)指标易受对抗攻击。攻击者可通过预处理篡改视频,使其在特定指标下得分虚高,但视觉质量无实质提升。现有研究多集中于白盒攻击,而针对VQA的黑盒攻击关注较少。此外,已有研究指出,为某一模型生成的对抗样本在不同模型间缺乏可迁移性。本文提出一种跨模态攻击方法IC2VQA,旨在揭示现代VQA模型的脆弱性。该方法基于图像与视频低层特征空间相似性的观察,探究跨模态对抗扰动的可迁移性;具体而言,分析在带额外CLIP模块的白盒IQA模型上生成的扰动,是否能有效攻击VQA模型。实验表明,加入CLIP模块显著提升了扰动的迁移能力,因CLIP能有效捕捉低层语义。大量实验显示,IC2VQA在三种黑盒VQA模型上均取得高攻击成功率。与现有黑盒攻击策略对比,本方法在相同迭代次数和攻击强度下表现更优。我们认为该方法有助于深入分析VQA指标的鲁棒性。
原文摘要 · Abstract (English)
Recent studies have revealed that modern image and video quality assessment (IQA/VQA) metrics are vulnerable to adversarial attacks. An attacker can manipulate a video through preprocessing to artificially increase its quality score according to a certain metric, despite no actual improvement in visual quality. Most of the attacks studied in the literature are white-box attacks, while black-box attacks in the context of VQA have received less attention. Moreover, some research indicates a lack of transferability of adversarial examples generated for one model to another when applied to VQA. In this paper, we propose a cross-modal attack method, IC2VQA, aimed at exploring the vulnerabilities of modern VQA models. This approach is motivated by the observation that the low-level feature spaces of images and videos are similar. We investigate the transferability of adversarial perturbations across different modalities; specifically, we analyze how adversarial perturbations generated on a white-box IQA model with an additional CLIP module can effectively target a VQA model. The addition of the CLIP module serves as a valuable aid in increasing transferability, as the CLIP model is known for its effective capture of low-level semantics. Extensive experiments demonstrate that IC2VQA achieves a high success rate in attacking three black-box VQA models. We compare our method with existing black-box attack strategies, highlighting its superiority in terms of attack success within the same number of iterations and levels of attack strength. We believe that the proposed method will contribute to the deeper analysis of robust VQA metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。