arXiv:2606.01481cs.CV2026-06被引 2

针对图文生成视频的安全风险,构建了首个综合评测基准。

SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation

论文配图:SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation
图 1 · 摘自论文原文
  • 设计10类恶意内容类别,评估图文联合输入下的生成安全
  • 现有模型在高画质下不安全得分高达44.5,防护能力弱
  • 单一文本或图像防护失效率达80%,需多模态协同防御

随着文本到图像扩散模型的快速发展,Sora等文本到视频生成模型(T2V)已能根据文本提示或初始图像生成短时合成视频。然而,尤其在图像引导下,合成视频常存在生成非法、政治敏感或不道德内容的风险。现有基准主要测试恶意文本提示下的安全性,忽视了文本与图像组合仍可能引发有害内容的场景。实际中,这一问题常见且棘手:即使输入的文本和图像本身安全,生成的视频仍可能传递有害信息。为此,我们提出SafeGen-Bench,一个专为评估条件化T2V模型安全性设计的基准。该基准涵盖来自多样化图像和视频源的起始帧,搭配对应文本提示,模拟真实输入场景。我们评估多种条件化T2V模型,结果显示当前模型难以一致避免生成恶意内容,不安全得分最高达44.5,尤其在高画质条件下更为明显。此外,我们测试了基于文本和图像的防护机制,发现单一模态防护在七个恶意类别中失败率高达80%,难以提供可靠防御。我们希望SafeGen-Bench能推动更安全、可控的条件化T2V模型发展。

原文摘要 · Abstract (English)

With the rapid advancements in text-to-image diffusion models, generative video models (T2V models) like Sora can now produce short synthetic videos from a text prompt or an initial image. However, synthetic video generation -- especially when guided by an initial image -- often poses risks, including the potential creation of illegal, politically sensitive, or unethical content. Existing benchmarks have started to consider the safety of generated videos, but they primarily focus on testing models with malicious text prompts, ignoring the scenario where text prompt and image combination may still lead to harmful video content. In practice, this is a common and challenging issue: videos generated from safe text and image inputs can nonetheless convey harmful information. To bridge this gap, we introduce SafeGen-Bench, a benchmark specifically designed to evaluate the safety of conditional T2V models. Our benchmark defines 10 malicious categories, concentrating on risks related to both temporal sequences and depicted behaviors. SafeGen-Bench consists of carefully selected start frames from diverse image and video sources, paired with corresponding text prompts to simulate realistic inputs. We evaluate a variety of conditional T2V models on SafeGen-Bench, and the results indicate that current models struggle to consistently avoid generating malicious content with unsafety scores reaching up to 44.5, especially under conditions requiring high quality. Furthermore, we assess the effectiveness of both text-based and image-based guardrails on our benchmark, finding that unimodal guardrails alone were insufficient to provide a robust defense, with an 80\% failure rate across seven malicious categories. We hope that SafeGen-Bench will foster the development of safer and more controllable conditional T2V models.

视频生成安全评测多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。