让自回归图像生成模型彻底删除特定概念,防止生成不当内容
Obliviate: Erasing Concepts from Autoregressive Image Generation Models

- 用视觉令牌分布的KL散度指导,精准定位需删除的概念
- 在完整生成流程中更新轨迹,实现稳定且高效的去除非必要内容
- 适用于需要安全可控生成的场景,如内容审核与合规应用
生成式AI的广泛应用加剧了滥用风险,例如生成不适宜或令人不适的图像。为缓解此类问题,已有若干概念擦除方法被提出以消除多模态生成模型中的有害内容。然而,自回归图像生成模型的概念擦除仍鲜有研究,尽管这类模型在统一多模态架构趋势中日益重要。本文提出Obliviate,一种基于引导的自回归图像生成概念擦除方法。其核心设计包括:基于KL的视觉令牌分布监督、对完整自回归推理轨迹的更新、以及用于稳定目标构建的对齐视觉前缀。我们在三个最先进的自回归文本到图像模型(Liquid、Emu3-Gen、Janus-Pro)上评估该方法,涵盖裸露内容、暴力画面及品牌图像的擦除。Obliviate持续优于现有方法,在防御性RAB基准上将裸露内容比例从91.58%降至3.15%,同时保持模型整体可用性。
原文摘要 · Abstract (English)
The widespread adoption of generative AI models has intensified concerns about misuse, including the creation of unsafe or disturbing imagery. To mitigate such issues, several concept erasure approaches have been proposed to remove harmful content from multimodal generative models. Yet concept erasure for autoregressive image generation remains largely unexplored, despite the growing relevance of these models in recent trends toward unified multimodal architectures. In this work, we fill this gap by introducing Obliviate, a guidance-based concept erasure method for autoregressive image generation. Our method builds on three key design choices: KL-based supervision over visual token distributions, trajectory-level updates over full autoregressive rollouts, and aligned visual prefixes for stable target construction. We evaluate Obliviate on three state-of-the-art autoregressive text-to-image models, Liquid, Emu3-Gen, and Janus-Pro, covering the erasure of explicit content, graphic violence, and branded imagery. Obliviate consistently outperforms current alternatives, reducing nudity on the defensive RAB benchmark from 91.58 to 3.15 while preserving overall model utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。