arXiv:2506.07891cs.CV2025-06被引 8

无需重训练即可清除视频生成模型中的有害概念。

Video Unlearning via Low-Rank Refusal Vector

  • 通过安全/不安全提示对估计拒绝向量,闭式更新模型权重。
  • 在多个基准上降低36.3%~58.2%的有害生成率。
  • 保持生成质量与提示对齐,适合安全可控视频生成场景。

视频生成模型通过利用大规模网络数据实现高质量文本到视频合成,但其训练范式易引入有害偏见与不良概念,导致生成非法或不当内容的风险。现有机器遗忘方法或依赖过滤(易被绕过),或需代价高昂的微调或无训练闭式修改。本文提出首个面向视频扩散模型的概念删除训练自由框架:基于五组安全/不安全提示对估计拒绝向量,并以闭式方式嵌入模型权重。通过对比低秩分解,有效解耦目标概念与无关语义,实现选择性抑制且不影响生成质量。该方法在Open-Sora和ZeroScopeT2V模型上,于T2VSafetyBench和SafeSora基准下分别实现平均36.3%和58.2%的有害生成减少,同时保持提示对齐与视频质量,为安全视频生成提供高效可扩展的解决方案,无需重训练或推理开销。

原文摘要 · Abstract (English)

Video generative models achieve high-quality synthesis from natural-language prompts by leveraging large-scale web data. However, this training paradigm inherently exposes them to unsafe biases and harmful concepts, introducing the risk of generating undesirable or illicit content. To mitigate unsafe generations, existing machine unlearning approaches either rely on filtering, and can therefore be bypassed, or they update model weights, but with costly fine-tuning or training-free closed-form edits. We propose the first training-free weight update framework for concept removal in video diffusion models. From five paired safe/unsafe prompts, our method estimates a refusal vector and integrates it into the model weights as a closed-form update. A contrastive low-rank factorization further disentangles the target concept from unrelated semantics, it ensures a selective concept suppression and it does not harm generation quality. Our approach reduces unsafe generations on the Open-Sora and ZeroScopeT2V models across the T2VSafetyBench and SafeSora benchmarks, with average reductions of 36.3% and 58.2% respectively, while preserving prompt alignment and video quality. This establishes an efficient and scalable solution for safe video generation without retraining nor any inference overhead. Project page: https://www.pinlab.org/video-unlearning.

视频生成扩散模型安全生成去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。