无需训练即可消除文本生成视频中的不良概念。
VideoEraser: Concept Erasure in Text-to-Video Diffusion Models
- 通过两阶段模块动态调整提示嵌入与噪声引导
- 四大任务平均降低46%的有害内容生成率
- 可直接集成到现有文生视频模型中,适合安全管控场景
文本生成视频(T2V)扩散模型的快速发展引发了隐私、版权和安全问题,因其可能被用于生成有害或误导性内容。这些模型常在包含未经授权的个人身份、艺术创作和有害材料的数据集上训练,导致此类内容的不可控生成与传播。为解决该问题,我们提出VideoEraser,一种无需训练的框架,可在明确提示时仍阻止T2V扩散模型生成特定不良概念的视频。该框架作为即插即用模块,通过两阶段流程——选择性提示嵌入调节(SPEA)与对抗鲁棒噪声引导(ARNG),无缝集成至主流T2V扩散模型。我们在物体擦除、艺术风格擦除、名人擦除和裸露内容擦除四项任务中进行广泛评估。实验表明,VideoEraser在有效性、完整性、保真度、鲁棒性和泛化能力方面均优于现有方法,尤其在抑制有害内容生成上达到当前最优性能,相较基线平均降低46%。
原文摘要 · Abstract (English)
The rapid growth of text-to-video (T2V) diffusion models has raised concerns about privacy, copyright, and safety due to their potential misuse in generating harmful or misleading content. These models are often trained on numerous datasets, including unauthorized personal identities, artistic creations, and harmful materials, which can lead to uncontrolled production and distribution of such content. To address this, we propose VideoEraser, a training-free framework that prevents T2V diffusion models from generating videos with undesirable concepts, even when explicitly prompted with those concepts. Designed as a plug-and-play module, VideoEraser can seamlessly integrate with representative T2V diffusion models via a two-stage process: Selective Prompt Embedding Adjustment (SPEA) and Adversarial-Resilient Noise Guidance (ARNG). We conduct extensive evaluations across four tasks, including object erasure, artistic style erasure, celebrity erasure, and explicit content erasure. Experimental results show that VideoEraser consistently outperforms prior methods regarding efficacy, integrity, fidelity, robustness, and generalizability. Notably, VideoEraser achieves state-of-the-art performance in suppressing undesirable content during T2V generation, reducing it by 46% on average across four tasks compared to baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。