首次提出针对文生视频检索的视频推广攻击,提升目标视频排名。
Adversarial Video Promotion Against Text-to-Video Retrieval
- 设计对抗性攻击,主动提升特定视频在多个查询下的排名。
- 在白/灰/黑盒场景下平均超越基线30%/10%/4%的推广效果。
- 适用于研究模型安全、评估多模态系统鲁棒性的研究人员。
随着跨模态模型的发展,文本到视频检索(T2VR)迅速进步,但其鲁棒性尚未得到充分研究。现有攻击主要旨在降低视频排名,而能将视频推向特定查询、提升其排名的攻击仍基本未被探索。此类攻击更具影响力,因攻击者可通过提高视频曝光量获取经济利益或传播误导信息。为此,我们首次提出针对T2VR的对抗性视频推广攻击(ViPro),并引入模态精炼(MoRe)机制,以捕捉视觉与文本模态间更细粒度的复杂交互,增强黑盒迁移能力。实验涵盖2个基线、3个主流T2VR模型、3个常用数据集(超过10,000个视频),在三种场景下进行多目标评估,模拟攻击者同时提升多个查询相关视频的真实场景。还评估了防御能力和不可察觉性。结果表明,ViPro在白/灰/黑盒设置下平均优于基线30%/10%/4%。本工作揭示了被忽视的安全漏洞,提供了攻击上限/下限的定性分析,并为应对策略提供启示。代码将公开于 https://github.com/michaeltian108/ViPro。
原文摘要 · Abstract (English)
Thanks to the development of cross-modal models, text-to-video retrieval (T2VR) is advancing rapidly, but its robustness remains largely unexamined. Existing attacks against T2VR are designed to push videos away from queries, i.e., suppressing the ranks of videos, while the attacks that pull videos towards selected queries, i.e., promoting the ranks of videos, remain largely unexplored. These attacks can be more impactful as attackers may gain more views/clicks for financial benefits and widespread (mis)information. To this end, we pioneer the first attack against T2VR to promote videos adversarially, dubbed the Video Promotion attack (ViPro). We further propose Modal Refinement (MoRe) to capture the finer-grained, intricate interaction between visual and textual modalities to enhance black-box transferability. Comprehensive experiments cover 2 existing baselines, 3 leading T2VR models, 3 prevailing datasets with over 10k videos, evaluated under 3 scenarios. All experiments are conducted in a multi-target setting to reflect realistic scenarios where attackers seek to promote the video regarding multiple queries simultaneously. We also evaluated our attacks for defences and imperceptibility. Overall, ViPro surpasses other baselines by over $30/10/4\%$ for white/grey/black-box settings on average. Our work highlights an overlooked vulnerability, provides a qualitative analysis on the upper/lower bound of our attacks, and offers insights into potential counterplays. Code will be publicly available at https://github.com/michaeltian108/ViPro.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。