发现文本到视频模型删了敏感内容仍可能残留,新方法能检测其重新出现的潜力。
PROBE: Diagnosing Residual Concept Capacity in Erased Text-to-Video Diffusion Models
- 通过优化伪标记嵌入,检测被删除概念在视频中的潜在复活能力。
- 三种模型、三类概念测试显示,所有擦除方法都留下可测量的残留能力。
- 发现视频特有故障:概念随帧数渐进重现,传统方法无法察觉。
文本到视频扩散模型的概念擦除技术虽宣称大幅抑制敏感内容,但现有评估仅检查生成帧中目标概念是否消失,将输出级抑制视为表征移除的证据。我们提出PROBE诊断协议,量化被擦除概念在T2V模型中的重激活潜力。在冻结所有模型参数的前提下,PROBE通过去噪重建目标,结合新颖的潜在对齐约束,将恢复锚定于原概念的时空结构。主要贡献包括:(1) 构建多层级评估框架,涵盖分类器检测、语义相似度、时间重激活分析与人工验证;(2) 在三种T2V架构、三类概念及三种擦除策略下系统实验,发现所有方法均留有可测量的残余容量,其鲁棒性与干预深度正相关;(3) 识别出时间再涌现这一视频特有失效模式——被压制的概念在帧序列中逐步复现,无法被帧级指标捕捉。结果表明当前擦除方法仅实现输出级抑制,而非表征级移除。我们开源该协议以支持可复现的安全审计,代码见https://github.com/YiweiXie/PRObingBasedEvaluation。
原文摘要 · Abstract (English)
Concept erasure techniques for text-to-video (T2V) diffusion models report substantial suppression of sensitive content, yet current evaluation is limited to checking whether the target concept is absent from generated frames, treating output-level suppression as evidence of representational removal. We introduce PROBE, a diagnostic protocol that quantifies the \textit{reactivation potential} of erased concepts in T2V models. With all model parameters frozen, PROBE optimizes a lightweight pseudo-token embedding through a denoising reconstruction objective combined with a novel latent alignment constraint that anchors recovery to the spatiotemporal structure of the original concept. We make three contributions: (1) a multi-level evaluation framework spanning classifier-based detection, semantic similarity, temporal reactivation analysis, and human validation; (2) systematic experiments across three T2V architectures, three concept categories, and three erasure strategies revealing that all tested methods leave measurable residual capacity whose robustness correlates with intervention depth; and (3) the identification of temporal re-emergence, a video-specific failure mode where suppressed concepts progressively resurface across frames, invisible to frame-level metrics. These findings suggest that current erasure methods achieve output-level suppression rather than representational removal. We release our protocol to support reproducible safety auditing. Our code is available at https://github.com/YiweiXie/PRObingBasedEvaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。