arXiv:2411.05544cs.CVcs.LG2024-11被引 2

解决文本生成图像模型终身学习中的遗忘问题

Towards Lifelong Few-Shot Customization of Text-to-Image Diffusion

  • 提出无数据知识蒸馏与上下文生成机制,避免旧知识丢失
  • 在仅用少量新数据情况下仍能保持原有生成能力
  • 适合需要持续更新图像模型的个性化应用

文本到图像扩散模型的终身少样本定制旨在以极少数据持续适应新任务的同时保留旧知识。现有方法在少样本任务中表现良好,但在长期生成中面临灾难性遗忘问题。本文将遗忘现象分为相关概念遗忘和旧概念遗忘两类。为此,提出一种无需真实数据或历史数据重放的无数据知识蒸馏策略,实现边学边保留旧知识。同时设计基于上下文生成(ICGen)范式,使模型可基于输入视觉上下文进行条件生成,提升少样本生成效果并缓解旧知识遗忘。大量实验表明,所提LFS-Diffusion方法在保持高质量、高准确度图像生成的同时,有效维持已有知识。

原文摘要 · Abstract (English)

Lifelong few-shot customization for text-to-image diffusion aims to continually generalize existing models for new tasks with minimal data while preserving old knowledge. Current customization diffusion models excel in few-shot tasks but struggle with catastrophic forgetting problems in lifelong generations. In this study, we identify and categorize the catastrophic forgetting problems into two folds: relevant concepts forgetting and previous concepts forgetting. To address these challenges, we first devise a data-free knowledge distillation strategy to tackle relevant concepts forgetting. Unlike existing methods that rely on additional real data or offline replay of original concept data, our approach enables on-the-fly knowledge distillation to retain the previous concepts while learning new ones, without accessing any previous data. Second, we develop an In-Context Generation (ICGen) paradigm that allows the diffusion model to be conditioned upon the input vision context, which facilitates the few-shot generation and mitigates the issue of previous concepts forgetting. Extensive experiments show that the proposed Lifelong Few-Shot Diffusion (LFS-Diffusion) method can produce high-quality and accurate images while maintaining previously learned knowledge.

扩散模型少样本学习终身学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。