提出多模态风险检测框架,提前识别图文生成视频中的安全隐患。
ConceptGuard: Proactive Safety in Text-and-Image-to-Video Generation through Multimodal Risk Detection
- 通过融合图文输入构建概念空间,对比检测潜在风险
- 在生成过程中干预提示,避免产生不安全内容
- 首个针对图文到视频生成的安全评测基准,适合安全研究者
近年来,视频生成模型已能根据结合文本和图像的多模态提示生成高质量视频。尽管这些系统提升了可控性,但也引入了新安全风险,有害内容可能源自单一模态或其交互。现有方法多为纯文本、需预先知晓风险类别,或仅作为生成后的审计工具,难以主动应对组合性多模态风险。为此,我们提出 ConceptGuard,一个统一的安全防护框架,可主动检测并缓解多模态视频生成中的不安全语义。该框架分两阶段运行:首先,对比检测模块将融合的图文输入投影至结构化概念空间,识别潜在安全风险;其次,语义抑制机制通过干预提示的多模态条件,引导生成过程避开不安全概念。为支持框架开发与严格评估,我们引入两个新基准:ConceptRisk——用于训练多模态风险的大规模数据集;T2VSafetyBench-TI2V——首个适配文本与图像到视频(TI2V)场景的基准,源自 T2VSafetyBench。在两个基准上的综合实验表明,ConceptGuard 持续优于现有基线,在风险检测与安全视频生成方面均达最先进水平。
原文摘要 · Abstract (English)
Recent progress in video generative models has enabled the creation of high-quality videos from multimodal prompts that combine text and images. While these systems offer enhanced controllability, they also introduce new safety risks, as harmful content can emerge from individual modalities or their interaction. Existing safety methods are often text-only, require prior knowledge of the risk category, or operate as post-generation auditors, struggling to proactively mitigate such compositional, multimodal risks. To address this challenge, we present ConceptGuard, a unified safeguard framework for proactively detecting and mitigating unsafe semantics in multimodal video generation. ConceptGuard operates in two stages: First, a contrastive detection module identifies latent safety risks by projecting fused image-text inputs into a structured concept space; Second, a semantic suppression mechanism steers the generative process away from unsafe concepts by intervening in the prompt's multimodal conditioning. To support the development and rigorous evaluation of this framework, we introduce two novel benchmarks: ConceptRisk, a large-scale dataset for training on multimodal risks, and T2VSafetyBench-TI2V, the first benchmark adapted from T2VSafetyBench for the Text-and-Image-to-Video (TI2V) safety setting. Comprehensive experiments on both benchmarks show that ConceptGuard consistently outperforms existing baselines, achieving state-of-the-art results in both risk detection and safe video generation. Our code is available at https://github.com/Ruize-Ma/ConceptGuard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。