提出六维评估框架,全面测试图像生成模型去记忆能力
Holistic Unlearning Benchmark: A Multi-Faceted Evaluation for Text-to-Image Diffusion Model Unlearning
- 构建六维度评估体系,覆盖准确性、一致性等关键指标
- 测试33个目标概念,每类1.6万条提示词,涵盖名人等四类敏感内容
- 开源数据集与代码,助力更可靠的隐私保护模型研发
随着文本到图像扩散模型在商业应用中日益普及,对其不当或有害使用的担忧也不断上升,包括未经授权生成受版权保护或敏感内容。概念去记忆作为一种有前景的解决方案,可通过移除预训练模型中的不良信息来应对这些挑战。然而,现有评估主要关注目标概念是否被清除且图像质量是否保持,忽略了其他潜在影响,如副作用。本文提出全面的去记忆评估基准(Holistic Unlearning Benchmark, HUB),涵盖六个关键维度:忠实性、对齐性、精准性、多语言鲁棒性、抗攻击性与效率。该基准包含33个目标概念,每个概念对应16,000条提示词,覆盖名人、风格、知识产权和NSFW四类内容。我们的研究发现,没有任何一种方法能在所有评价标准上表现最优。通过公开评估代码与数据集,我们希望推动该领域进一步研究,发展出更可靠、高效的去记忆方法。
原文摘要 · Abstract (English)
As text-to-image diffusion models gain widespread commercial applications, there are increasing concerns about unethical or harmful use, including the unauthorized generation of copyrighted or sensitive content. Concept unlearning has emerged as a promising solution to these challenges by removing undesired and harmful information from the pre-trained model. However, the previous evaluations primarily focus on whether target concepts are removed while preserving image quality, neglecting the broader impacts such as unintended side effects. In this work, we propose Holistic Unlearning Benchmark (HUB), a comprehensive framework for evaluating unlearning methods across six key dimensions: faithfulness, alignment, pinpoint-ness, multilingual robustness, attack robustness, and efficiency. Our benchmark covers 33 target concepts, including 16,000 prompts per concept, spanning four categories: Celebrity, Style, Intellectual Property, and NSFW. Our investigation reveals that no single method excels across all evaluation criteria. By releasing our evaluation code and dataset, we hope to inspire further research in this area, leading to more reliable and effective unlearning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。