提出首个针对视觉自回归模型的概念擦除方法,解决生成内容安全问题。
SACE: Concept Erasure at the Semantic Singularity in Visual Autoregressive Models

- 基于语义奇点理论,仅在第一尺度干预实现精准擦除
- 在多个领域实现手术级概念删除,训练开销极低
- 适合关注生成内容安全的AI研究者与开发者
视觉自回归(VAR)模型在高保真文生图合成中取得突破,但生成内容的安全对齐问题日益突出。现有擦除技术直接应用于VAR模型会导致语义崩溃和视觉伪影,因其主要针对扩散模型的同质去噪步骤设计。为此,我们首次提出语义奇点公理:提示词中的目标语义概念必然锁定在Scale-0。通过增量语义显著性分析(ISSA)严格验证该公理,并实现粗到细语义注入过程的透明可查。基于此洞察,我们提出首个面向VAR模型的尺度感知概念擦除框架SACE。通过将干预严格限制在第一尺度,结合熵正则化擦除目标防止高熵采样退化,以及恢复性保留损失安全锚定纠缠良性先验。大量实验表明,本方法在多个领域实现精确概念擦除,且训练开销极小,有效解决新兴VAR架构固有的关键安全漏洞。代码已开源。
原文摘要 · Abstract (English)
The rapid progress of visual autoregressive (VAR) models has unlocked a transformative frontier for high-fidelity text-to-image synthesis, while heightening concerns over the safety alignment of generated content. Naive application of existing erasure techniques to VAR models causes catastrophic semantic collapse and visual artifacts, since they are predominantly designed for the homogeneous denoising steps of diffusion models. To address this foundational challenge, we first propose the Semantic Singularity Axiom, which posits that any target semantic concept embedded within a prompt is definitively locked at Scale-0. Then rigorously validate this axiom through our proposed Incremental Semantic Saliency Analysis (ISSA),which also enable the community to transparently inspect the coarse-to-fine semantic injection process. Guided by this insight, we introduce the first scale-aware concept erasure framework (SACE) for VAR models. By strictly confining interventions to the first scale, our approach couples an Entropy-Regularized Erasure Objective to prevent high-entropy sampling degeneration, alongside a restorative preservation loss to safely anchor the integrity of entangled benign priors. Extensive experiments demonstrate that our method achieves surgical concept erasure performance across various domains with minimal training overhead, timely and elegently resolute the critical safety vulnerabilities inherent in emerging VAR architectures. Code is available at: https://github.com/limerenceysy/SACE}{https://github.com/limerenceysy/SACE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。