arXiv:2410.11686cs.CV2024-10综述被引 1

用核理论统一梳理低样本视觉语言模型适配方法

A Survey of Low-shot Vision-Language Model Adaptation via Representer Theorem

  • 基于表示定理构建统一计算框架,解析现有方法本质
  • 在11个数据集上验证新方法显著提升低样本性能
  • 适合研究模型高效微调与理论分析的学者参考

预训练视觉-语言基础模型已推动零/少样本图像识别的发展。在训练数据有限的情况下,如何参数高效地微调这些模型成为关键挑战。此前已有大量相关方法提出,少数综述也进行了总结,但缺乏统一的计算框架来整合各类方法、揭示其本质并支持深入比较。本文首次从表示定理视角提出统一框架,并通过特化该框架推导出多种现有方法。进一步开展对比分析,揭示方法间的差异与联系。基于分析结果,提出若干改进方向。以在再生核希尔伯特空间(RKHS)中建模表示之间的类间相关性为例,利用核岭回归的闭式解实现扩展。在11个数据集上进行大量实验,验证了该方法的有效性。最后讨论局限性并提出未来研究方向。

原文摘要 · Abstract (English)

The advent of pre-trained vision-language foundation models has revolutionized the field of zero/few-shot (i.e., low-shot) image recognition. The key challenge to address under the condition of limited training data is how to fine-tune pre-trained vision-language models in a parameter-efficient manner. Previously, numerous approaches tackling this challenge have been proposed. Meantime, a few survey papers are also published to summarize these works. However, there still lacks a unified computational framework to integrate existing methods together, identify their nature and support in-depth comparison. As such, this survey paper first proposes a unified computational framework from the perspective of Representer Theorem and then derives many of the existing methods by specializing this framework. Thereafter, a comparative analysis is conducted to uncover the differences and relationships between existing methods. Based on the analyses, some possible variants to improve the existing works are presented. As a demonstration, we extend existing methods by modeling inter-class correlation between representers in reproducing kernel Hilbert space (RKHS), which is implemented by exploiting the closed-form solution of kernel ridge regression. Extensive experiments on 11 datasets are conducted to validate the effectiveness of this method. Toward the end of this paper, we discuss the limitations and provide further research directions.

视觉语言低样本学习表示定理综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。