arXiv:2602.10143cs.CV2026-02AAAI

用多模态增强原型,让少样本学习更准更快

MPA: Multimodal Prototype Augmentation for Few-Shot Learning

  • 用大模型生成多样化语义描述,丰富支持集信息
  • 结合多视角和自然增强,提升特征多样性
  • 通过不确定性建模吸收模糊样本,适合医学图像等场景

少样本学习(FSL)旨在仅凭少量标注样本识别新类别,广泛应用于自然科学、遥感和医学图像等领域。然而,现有方法多局限于视觉模态,直接从原始支持图像计算原型,缺乏丰富的多模态信息。为此,我们提出新型多模态原型增强框架MPA,包含基于大语言模型的多变语义增强(LMSE)、分层多视图增强(HMA)和自适应不确定类吸收器(AUCA)。LMSE利用大语言模型生成多样化的类别描述,为支持集注入额外语义线索;HMA结合自然与多视角增强(如视角距离、相机角度、光照变化)提升特征多样性;AUCA通过插值和高斯采样引入不确定类,有效吸收不确定样本。在四个单域和六个跨域FSL基准上的大量实验表明,MPA在多数设置下优于现有最优方法。特别地,在5类1样本设置下,单域和跨域场景分别超越第二好方法12.29%和24.56%。

原文摘要 · Abstract (English)

Recently, few-shot learning (FSL) has become a popular task that aims to recognize new classes from only a few labeled examples and has been widely applied in fields such as natural science, remote sensing, and medical images. However, most existing methods focus only on the visual modality and compute prototypes directly from raw support images, which lack comprehensive and rich multimodal information. To address these limitations, we propose a novel Multimodal Prototype Augmentation FSL framework called MPA, including LLM-based Multi-Variant Semantic Enhancement (LMSE), Hierarchical Multi-View Augmentation (HMA), and an Adaptive Uncertain Class Absorber (AUCA). LMSE leverages large language models to generate diverse paraphrased category descriptions, enriching the support set with additional semantic cues. HMA exploits both natural and multi-view augmentations to enhance feature diversity (e.g., changes in viewing distance, camera angles, and lighting conditions). AUCA models uncertainty by introducing uncertain classes via interpolation and Gaussian sampling, effectively absorbing uncertain samples. Extensive experiments on four single-domain and six cross-domain FSL benchmarks demonstrate that MPA achieves superior performance compared to existing state-of-the-art methods across most settings. Notably, MPA surpasses the second-best method by 12.29% and 24.56% in the single-domain and cross-domain setting, respectively, in the 5-way 1-shot setting.

少样本学习多模态原型增强大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。