arXiv:2509.22400cs.CV2025-09中稿 · ICLR被引 7

针对视觉自回归模型的安全漏洞,提出精准擦除不安全概念的新方法。

Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models

  • 用辅助视觉标记降低微调强度,实现稳定概念擦除。
  • 通过过滤交叉熵损失精确定位并最小化调整危险视觉标记。
  • 保持生成质量与语义一致性,适合需高安全性图像生成的场景。

视觉自回归(VAR)模型在文生图领域进展迅速,但安全问题日益突出。现有概念擦除技术多针对扩散模型,难以适配VAR模型的逐标记预测机制。本文提出VARE框架,利用辅助视觉标记降低微调强度,实现稳定擦除。在此基础上,提出S-VARE方法,引入过滤交叉熵损失以精准识别并最小化调整不安全视觉标记,并结合保留损失维持语义保真度,解决粗略微调带来的语言漂移与多样性下降问题。大量实验表明,该方法可在保持生成质量的同时实现精准概念擦除,有效弥补此前方法在自回归文生图中的安全缺口。

原文摘要 · Abstract (English)

The rapid progress of visual autoregressive (VAR) models has brought new opportunities for text-to-image generation, but also heightened safety concerns. Existing concept erasure techniques, primarily designed for diffusion models, fail to generalize to VARs due to their next-scale token prediction paradigm. In this paper, we first propose a novel VAR Erasure framework VARE that enables stable concept erasure in VAR models by leveraging auxiliary visual tokens to reduce fine-tuning intensity. Building upon this, we introduce S-VARE, a novel and effective concept erasure method designed for VAR, which incorporates a filtered cross entropy loss to precisely identify and minimally adjust unsafe visual tokens, along with a preservation loss to maintain semantic fidelity, addressing the issues such as language drift and reduced diversity introduce by naïve fine-tuning. Extensive experiments demonstrate that our approach achieves surgical concept erasure while preserving generation quality, thereby closing the safety gap in autoregressive text-to-image generation by earlier methods.

图像安全概念擦除自回归模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。