arXiv:2507.13739cs.CVcs.AI2025-07ICCV被引 5

用生成模型缓解少样本增量学习中的遗忘问题

Can Synthetic Images Conquer Forgetting? Beyond Unexplored Doubts in Few-Shot Class-Incremental Learning

  • 用冻结的文生图模型提取多尺度特征作为潜在记忆
  • 在多个数据集上显著降低遗忘率,新类识别准确率超现有方法
  • 适合研究生成模型与持续学习结合的学者参考

少样本增量学习(FSCIL)因训练数据极度有限而面临灾难性遗忘挑战。本文提出Diffusion-FSCIL,采用冻结的文生图扩散模型作为主干网络,利用其大规模预训练带来的生成能力、多尺度表征和文本编码器的表征灵活性。为最大化表征能力,通过提取多种互补的扩散特征作为潜在回放,并辅以轻量特征蒸馏防止生成偏差。该框架通过冻结主干、极少可训练参数及批量特征提取实现高效。在CUB-200、miniImageNet和CIFAR-100上的大量实验表明,该方法在保持旧类别性能的同时,有效适应新类别,优于当前最优方法。

原文摘要 · Abstract (English)

Few-shot class-incremental learning (FSCIL) is challenging due to extremely limited training data; while aiming to reduce catastrophic forgetting and learn new information. We propose Diffusion-FSCIL, a novel approach that employs a text-to-image diffusion model as a frozen backbone. Our conjecture is that FSCIL can be tackled using a large generative model's capabilities benefiting from 1) generation ability via large-scale pre-training; 2) multi-scale representation; 3) representational flexibility through the text encoder. To maximize the representation capability, we propose to extract multiple complementary diffusion features to play roles as latent replay with slight support from feature distillation for preventing generative biases. Our framework realizes efficiency through 1) using a frozen backbone; 2) minimal trainable components; 3) batch processing of multiple feature extractions. Extensive experiments on CUB-200, \emph{mini}ImageNet, and CIFAR-100 show that Diffusion-FSCIL surpasses state-of-the-art methods, preserving performance on previously learned classes and adapting effectively to new ones.

少样本学习增量学习生成模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。