用多视角协作优化提升视觉语言模型在跨域少样本学习中的表现
Cross-Domain Few-Shot Learning via Multi-View Collaborative Optimization with Vision-Language Models
- 设计双专家模块提取多视角特征,结合一致性约束增强鲁棒性
- 在跨域少样本任务中性能超越现有方法,新基准上显著领先
- 适合关注跨域迁移与少样本学习的研究者和应用开发者
基于自然图像与语言数据预训练的视觉-语言模型(如CLIP)在少样本图像识别中展现出巨大潜力,催生了多种高效的迁移学习方法。这些方法利用模型内在先验知识,在标准图像数据集上取得优异表现。然而,当面对成像域不同于自然图像的跨域任务时,其效果往往受限。为此,本文提出一致性引导的多视角协同优化(CoMuCo)策略,通过两个功能互补的专家模块提取多视角特征,并引入基于先验知识的一致性约束与信息几何共识机制,提升特征学习的鲁棒性。此外,构建了一个新的跨域少样本学习基准,以全面评估方法在非自然成像域的表现。在既有与新提出的基准上的大量实验表明,CoMuCo在少样本任务中持续优于当前主流方法。代码与基准将公开发布。
原文摘要 · Abstract (English)
Vision-language models (VLMs) pre-trained on natural image and language data, such as CLIP, have exhibited significant potential in few-shot image recognition tasks, leading to development of various efficient transfer learning methods. These methods exploit inherent pre-learned knowledge in VLMs and have achieved strong performance on standard image datasets. However, their effectiveness is often limited when confronted with cross-domain tasks where imaging domains differ from natural images. To address this limitation, we propose Consistency-guided Multi-view Collaborative Optimization (CoMuCo), a novel fine-tuning strategy for VLMs. This strategy employs two functionally complementary expert modules to extract multi-view features, while incorporating prior knowledge-based consistency constraints and information geometry-based consensus mechanisms to enhance the robustness of feature learning. Additionally, a new cross-domain few-shot benchmark is established to help comprehensively evaluate methods on imaging domains distinct from natural images. Extensive empirical evaluations on both existing and newly proposed benchmarks suggest CoMuCo consistently outperforms current methods in few-shot tasks. The code and benchmark will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。