arXiv:2511.21145cs.CV2025-11被引 2

针对文本生成视频的动态安全风险,提出自动红队测试框架TEAR。

TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Models

  • 基于两阶段优化生成时间敏感提示,诱使模型输出违规视频。
  • 在多个开源与商用系统中攻击成功率超80%,优于此前57%最佳水平。
  • 适合关注视频生成安全、对抗样本研究的研究者使用。

文本到视频(T2V)模型能够生成高质量且时间连贯的动态视频内容,但其多样化的生成能力也带来了严峻的安全挑战。现有安全评估方法主要针对静态图像和文本生成,难以捕捉视频生成中的复杂时序动态。为此,我们提出一种时间感知自动化红队框架TEAR,专门用于发现与T2V模型动态时序特性相关联的安全风险。TEAR采用两阶段优化的时间感知测试生成器:先进行初始生成器训练,再通过在线偏好学习实现时间感知优化,以构造看似无害的文本提示,利用时序动态诱导模型产生违反政策的视频输出。同时引入精炼模型,循环提升提示的隐蔽性与对抗效果。大量实验表明,TEAR在开源与商业T2V系统中均表现出色,攻击成功率超过80%,较之前最佳结果(57%)有显著提升。

原文摘要 · Abstract (English)

Text-to-Video (T2V) models are capable of synthesizing high-quality, temporally coherent dynamic video content, but the diverse generation also inherently introduces critical safety challenges. Existing safety evaluation methods,which focus on static image and text generation, are insufficient to capture the complex temporal dynamics in video generation. To address this, we propose a TEmporal-aware Automated Red-teaming framework, named TEAR, an automated framework designed to uncover safety risks specifically linked to the dynamic temporal sequencing of T2V models. TEAR employs a temporal-aware test generator optimized via a two-stage approach: initial generator training and temporal-aware online preference learning, to craft textually innocuous prompts that exploit temporal dynamics to elicit policy-violating video output. And a refine model is adopted to improve the prompt stealthiness and adversarial effectiveness cyclically. Extensive experimental evaluation demonstrates the effectiveness of TEAR across open-source and commercial T2V systems with over 80% attack success rate, a significant boost from prior best result of 57%.

视频生成安全评估对抗攻击时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。