研究去学习对图像生成能力的副作用,发现删掉内容会破坏生成逻辑。
Erasure or Erosion? Evaluating Compositional Degradation in Unlearned Text-To-Image Diffusion Models
- 从组合生成角度评估文本到图像模型的去学习效果。
- 强去除效果常导致属性绑定、空间推理和计数能力显著下降。
- 适合关注模型安全与生成质量平衡的研究者。
事后去学习已成为从大型文本到图像扩散模型中移除不良概念的实用方法。然而,以往工作主要通过删除成功率评估去学习效果,对其对整体生成能力的影响了解甚少。本文针对 Stable Diffusion 1.4 中的裸露内容移除,系统性地评估了多种前沿去学习方法在 T2I-CompBench++ 与 GenEval 上的表现,同时结合传统去学习基准。结果揭示出一种普遍权衡:实现强删除效果的方法往往导致属性绑定、空间推理和计数能力严重退化;而保持组合结构完整的方法则难以提供稳健的删除效果。该发现凸显了现有评估方式的局限性,强调需设计兼顾语义保留与目标抑制的去学习目标。
原文摘要 · Abstract (English)
Post-hoc unlearning has emerged as a practical mechanism for removing undesirable concepts from large text-to-image diffusion models. However, prior work primarily evaluates unlearning through erasure success; its impact on broader generative capabilities remains poorly understood. In this work, we conduct a systematic empirical study of concept unlearning through the lens of compositional text-to-image generation. Focusing on nudity removal in Stable Diffusion 1.4, we evaluate a diverse set of state-of-the-art unlearning methods using T2I-CompBench++ and GenEval, alongside established unlearning benchmarks. Our results reveal a consistent trade-off between unlearning effectiveness and compositional integrity: methods that achieve strong erasure frequently incur substantial degradation in attribute binding, spatial reasoning, and counting. Conversely, approaches that preserve compositional structure often fail to provide robust erasure. These findings highlight limitations of current evaluation practices and underscore the need for unlearning objectives that explicitly account for semantic preservation beyond targeted suppression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。