arXiv:2504.16907cs.CVcs.AI2025-04ICCV被引 10

利用文本生成视频的冗余信息,实现隐蔽后门攻击。

BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation

论文配图:BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
图 1 · 摘自论文原文
  • 通过时空组合与动态元素变换,将恶意内容嵌入视频生成过程。
  • 攻击成功率高,且不破坏原始语义,可绕过传统内容审核。
  • 首次针对文生视频模型提出隐蔽后门攻击,适合安全研究者关注。

文本到视频(T2V)生成模型快速发展,广泛应用于娱乐、教育和营销等领域。然而,这类模型的对抗脆弱性尚未被充分研究。我们发现,在T2V生成任务中,生成视频常包含大量未在文本提示中明确指定的冗余信息,如环境元素、次要物体和附加细节,为恶意攻击者嵌入隐藏有害内容提供了机会。基于这一内在冗余,我们提出BadVideo——首个专用于T2V生成的后门攻击框架。攻击通过两种关键策略实现:(1) 时空组合,将不同时空特征融合以编码恶意信息;(2) 动态元素变换,在时间维度上对冗余元素进行变化以传递恶意信号。攻击目标能无缝融入用户文本指令,具备高度隐蔽性。此外,利用视频的时间维度特性,该攻击可有效规避主要依赖单帧空间信息的传统内容审核系统。大量实验表明,BadVideo在保持原始语义和清洁输入性能的同时,实现了高攻击成功率。本工作揭示了T2V模型的对抗脆弱性,警示潜在风险与滥用可能。项目页面见 https://wrt2000.github.io/BadVideo2025/。

原文摘要 · Abstract (English)

Text-to-video (T2V) generative models have rapidly advanced and found widespread applications across fields like entertainment, education, and marketing. However, the adversarial vulnerabilities of these models remain rarely explored. We observe that in T2V generation tasks, the generated videos often contain substantial redundant information not explicitly specified in the text prompts, such as environmental elements, secondary objects, and additional details, providing opportunities for malicious attackers to embed hidden harmful content. Exploiting this inherent redundancy, we introduce BadVideo, the first backdoor attack framework tailored for T2V generation. Our attack focuses on designing target adversarial outputs through two key strategies: (1) Spatio-Temporal Composition, which combines different spatiotemporal features to encode malicious information; (2) Dynamic Element Transformation, which introduces transformations in redundant elements over time to convey malicious information. Based on these strategies, the attacker's malicious target seamlessly integrates with the user's textual instructions, providing high stealthiness. Moreover, by exploiting the temporal dimension of videos, our attack successfully evades traditional content moderation systems that primarily analyze spatial information within individual frames. Extensive experiments demonstrate that BadVideo achieves high attack success rates while preserving original semantics and maintaining excellent performance on clean inputs. Overall, our work reveals the adversarial vulnerability of T2V models, calling attention to potential risks and misuse. Our project page is at https://wrt2000.github.io/BadVideo2025/.

后门攻击文生视频安全漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。