让模型像人一样持续学习新组合,提升视觉理解泛化能力。
Composition-Incremental Learning for Compositional Generalization
- 用视觉合成+语言表征蒸馏,实现增量学习新组合。
- 在MIT-States-CompIL和C-GQA-CompIL上显著提升泛化性能。
- 适合需要持续学习新场景的视觉系统研发者。
组合泛化在预收集训练数据的计算机视觉中已取得显著进展。然而,现实世界数据持续涌现,可能的组合近乎无限、长尾分布且不完全可见。因此,理想模型应能以增量方式逐步提升组合泛化能力。本文探索了在组合零样本学习(CZSL)任务下的组合增量学习(CompIL),使模型能持续学习新组合,逐步增强组合泛化能力。为定量评估CompIL,我们利用现有数据集构建了基准测试流程,生成MIT-States-CompIL和C-GQA-CompIL两个数据集。同时提出一种伪回放框架,结合视觉合成器生成已学组合的视觉表示,并通过语言基元蒸馏机制保持学习过程中的表征对齐。大量实验验证了该框架的有效性。
原文摘要 · Abstract (English)
Compositional generalization has achieved substantial progress in computer vision on pre-collected training data. Nonetheless, real-world data continually emerges, with possible compositions being nearly infinite, long-tailed, and not entirely visible. Thus, an ideal model is supposed to gradually improve the capability of compositional generalization in an incremental manner. In this paper, we explore Composition-Incremental Learning for Compositional Generalization (CompIL) in the context of the compositional zero-shot learning (CZSL) task, where models need to continually learn new compositions, intending to improve their compositional generalization capability progressively. To quantitatively evaluate CompIL, we develop a benchmark construction pipeline leveraging existing datasets, yielding MIT-States-CompIL and C-GQA-CompIL. Furthermore, we propose a pseudo-replay framework utilizing a visual synthesizer to synthesize visual representations of learned compositions and a linguistic primitive distillation mechanism to maintain aligned primitive representations across the learning process. Extensive experiments demonstrate the effectiveness of the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。