让多个异构模型协作,提升CLIP在跨域任务中的泛化能力。
TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaboration
- 通过统一框架整合11个异构专家模型的知识
- 低样本设置下平均比CoOp高10%,EuroSAT上高20%
- 无需额外计算开销,适合跨域视觉识别任务
视觉-语言基础模型(如CLIP)凭借大规模图文预训练展现出强大迁移能力。然而,下游任务的目标域数据与预训练阶段差异显著,导致单一模型泛化困难。尽管存在大量在不同模态、任务、网络和数据集上预训练的专家模型,但它们结构异构,彼此孤立,如何整合其知识以增强CLIP类模型仍不明确。为此,我们提出通用简洁的TransAgent框架,以统一方式传输孤立代理的知识,并通过多源知识蒸馏有效引导CLIP泛化。该框架灵活协同11个异构代理,不增加推理成本。最终,TransAgent在11个视觉识别数据集上达到当前最优性能,在相同低样本设置下,平均比CoOp高出约10%,在包含大领域偏移的EuroSAT上提升达20%。
原文摘要 · Abstract (English)
Vision-language foundation models (such as CLIP) have recently shown their power in transfer learning, owing to large-scale image-text pre-training. However, target domain data in the downstream tasks can be highly different from the pre-training phase, which makes it hard for such a single model to generalize well. Alternatively, there exists a wide range of expert models that contain diversified vision and/or language knowledge pre-trained on different modalities, tasks, networks, and datasets. Unfortunately, these models are "isolated agents" with heterogeneous structures, and how to integrate their knowledge for generalizing CLIP-like models has not been fully explored. To bridge this gap, we propose a general and concise TransAgent framework, which transports the knowledge of the isolated agents in a unified manner, and effectively guides CLIP to generalize with multi-source knowledge distillation. With such a distinct framework, we flexibly collaborate with 11 heterogeneous agents to empower vision-language foundation models, without further cost in the inference phase. Finally, our TransAgent achieves state-of-the-art performance on 11 visual recognition datasets. Under the same low-shot setting, it outperforms the popular CoOp with around 10% on average, and 20% on EuroSAT which contains large domain shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。