提出无需训练的视频生成概念删除方法,可精准擦除目标内容且保持视频质量。
Inference-Time Concept Suppression and Video-Centric Evaluation for Text-to-Video Models

- 推理时通过定位提示词证据抑制目标概念表达,不修改模型参数。
- 在五个概念上平均删除成功率70.4%,帧级残留率25.7%,优于现有方法。
- 专为视频设计评估体系,兼顾遗忘效果、非目标保留与生成质量。
文本到视频(T2V)生成器能合成真实且时间连贯的视频,但可控地从生成器中移除特定概念仍具挑战。与文本到图像的概念擦除不同,T2V 的去学习需在多帧中抑制目标概念,同时保留非目标主体、动作、场景及时间结构。本文提出 extbf{SIRUS},一种无需训练的推理时概念级 T2V 去学习框架。给定目标概念的文本别名,SIRUS 定位相关提示词证据并抑制采样过程中的目标表达,无需更新文本编码器或去噪网络。我们还引入面向视频的评估框架,分别衡量目标遗忘、非目标保留、视频质量、越狱鲁棒性与效率,采用视频级失败标准、帧级残留统计、成对保留分析、基于 VBench 的质量诊断及部署开销测量。在 CogVideoX 上对五类安全、物体与风格概念的实验显示,SIRUS 平均遗忘成功率达 70.4%,平均帧级命中率 25.7%,优于 VideoEraser 的 44.4% / 47.2%;同时将平均 VBench 质量下降从 -0.043 降低至 -0.016,实现最优遗忘-质量权衡。跨 Wan2.2 模型的迁移实验表明,SIRUS 可泛化至现代 T2V 主干网络。
原文摘要 · Abstract (English)
Text-to-video (T2V) generators can synthesize realistic and temporally coherent videos, but controllably removing a target concept from a generator remains difficult. Unlike text-to-image concept erasure, T2V unlearning must suppress a target concept that may persist across frames while preserving non-target subjects, actions, scenes, and temporal structure. We propose \textbf{SIRUS}, a training-free inference-time framework for concept-level T2V unlearning. Given textual aliases of a target concept, SIRUS localizes target-related prompt evidence and suppresses target expression during sampling, without updating the text encoder or denoising network. We further introduce a video-oriented evaluation framework for T2V unlearning that separately measures target forgetting, non-target preservation, video quality, jailbreak robustness, and efficiency, using video-level failure criteria, frame-level residue statistics, paired preservation analysis, VBench-based quality diagnostics, and deployment overhead measurement. Across five safety, object, and style concepts on CogVideoX, SIRUS reaches 70.4\% average forgetting success and 25.7\% average frame hit, compared with 44.4\% / 47.2\% for VideoEraser, while reducing the average VBench quality drop from -0.043 to -0.016, yielding the strongest forgetting-quality trade-off among fully evaluated baselines. Transfer experiments on Wan2.2 further suggest that SIRUS generalizes across modern T2V backbones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。