arXiv:2410.10663cs.CVcs.LG2024-10被引 1

跨模态少样本学习新框架,让模型像人一样抽象通用概念。

Cross-Modal Few-Shot Learning: a Generative Transfer Learning Framework

  • 构建生成式迁移框架,联合学习跨模态共享概念与模态内扰动。
  • 在7个跨模态数据集上达到当前最优性能,覆盖RGB-Sketch等三类模态组合。
  • 适合研究多模态少样本识别、知识迁移与人类认知启发模型的学者。

现有少样本学习研究主要聚焦单模态场景,即利用单一模态的少量标注数据泛化到未见数据。然而真实世界数据具有天然多模态特性,单模态方法限制了少样本学习的实际应用。为此,本文提出跨模态少样本学习(CFSL)任务,旨在利用稀缺标注数据实现跨多种模态的实例识别。该任务面临独特挑战,源于各模态间固有的视觉属性差异和结构不一致。为此,我们提出生成式迁移学习(GTL)框架,模拟人类抽象与泛化概念的方式。GTL通过生成结构联合估计跨模态潜在共享概念与模态内扰动。基于丰富单模态数据中潜在概念与视觉内容的关系,使模型能有效将知识从单模态迁移至新型多模态数据,如人类般理解。大量实验表明,GTL在七个跨模态数据集(涵盖RGB-Sketch、RGB-红外、RGB-深度)上均取得当前最优表现。

原文摘要 · Abstract (English)

Most existing studies on few-shot learning focus on unimodal settings, where models are trained to generalize to unseen data using a limited amount of labeled examples from a single modality. However, real-world data are inherently multi-modal, and such unimodal approaches limit the practical applications of few-shot learning. To bridge this gap, this paper introduces the Cross-modal Few-Shot Learning (CFSL) task, which aims to recognize instances across multiple modalities while relying on scarce labeled data. This task presents unique challenges compared to classical few-shot learning arising from the distinct visual attributes and structural disparities inherent to each modality. To tackle these challenges, we propose a Generative Transfer Learning (GTL) framework by simulating how humans abstract and generalize concepts. Specifically, the GTL jointly estimates the latent shared concept across modalities and the in-modality disturbance through a generative structure. Establishing the relationship between latent concepts and visual content among abundant unimodal data enables GTL to effectively transfer knowledge from unimodal to novel multimodal data, as humans did. Comprehensive experiments demonstrate that the GTL achieves state-of-the-art performance across seven multi-modal datasets across RGB-Sketch, RGB-Infrared, and RGB-Depth.

少样本学习跨模态生成模型知识迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。