arXiv:2409.17928cs.CLcs.AI2024-09EMNLP

构建细粒度数据集与新评估准则,提升文本到图像模型知识编辑的可靠性。

Pioneering Reliable Assessment in Text-to-Image Knowledge Editing: Leveraging a Fine-Grained Dataset and an Innovative Criterion

  • 设计细粒度数据集CAKE,支持多对象和改写测试,精准评估知识泛化。
  • 提出自适应CLIP阈值,有效剔除误判图像,实现更可靠的编辑评估。
  • 提出MPE方法,仅修改提示词即可更新知识,简单高效且优于现有模型编辑器。

在预训练阶段,文本到图像(T2I)扩散模型将事实知识编码于参数中,这些参数化知识虽能生成真实图像,但可能随时间过时,导致对世界状态的错误表征。知识编辑技术旨在靶向更新模型知识。然而,受限于编辑数据集不足和评估标准不可靠,当前T2I知识编辑难以有效泛化。本文提出一个三阶段框架:首先,构建名为CAKE的数据集,包含改写与多对象测试样本,支持更精细的知识泛化评估;其次,提出一种新型评估准则——自适应CLIP阈值,可有效过滤当前标准下误判的成功图像,实现可靠评估;最后,引入简单高效的编辑方法MPE,不依赖参数调优,而是精准识别并修改条件提示词中的过时部分,以适配最新知识。基于上下文学习的MPE实现版本,在整体性能上优于此前模型编辑器。本工作期望推动T2I知识编辑的可信评估发展。

原文摘要 · Abstract (English)

During pre-training, the Text-to-Image (T2I) diffusion models encode factual knowledge into their parameters. These parameterized facts enable realistic image generation, but they may become obsolete over time, thereby misrepresenting the current state of the world. Knowledge editing techniques aim to update model knowledge in a targeted way. However, facing the dual challenges posed by inadequate editing datasets and unreliable evaluation criterion, the development of T2I knowledge editing encounter difficulties in effectively generalizing injected knowledge. In this work, we design a T2I knowledge editing framework by comprehensively spanning on three phases: First, we curate a dataset \textbf{CAKE}, comprising paraphrase and multi-object test, to enable more fine-grained assessment on knowledge generalization. Second, we propose a novel criterion, \textbf{adaptive CLIP threshold}, to effectively filter out false successful images under the current criterion and achieve reliable editing evaluation. Finally, we introduce \textbf{MPE}, a simple but effective approach for T2I knowledge editing. Instead of tuning parameters, MPE precisely recognizes and edits the outdated part of the conditioning text-prompt to accommodate the up-to-date knowledge. A straightforward implementation of MPE (Based on in-context learning) exhibits better overall performance than previous model editors. We hope these efforts can further promote faithful evaluation of T2I knowledge editing methods.

知识编辑图像生成评估准则扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。