用二阶微分方程提升跨模态少样本学习效果
Cross-Modal Few-Shot Learning with Second-Order Neural Ordinary Differential Equations
- 用二阶神经微分方程建模特征演化,增强表达能力
- 在多个数据集上超越现有最优方法,显著减少过拟合
- 适合需要小样本下跨模态泛化的研究者使用
我们提出 SONO,一种基于二阶神经常微分方程(Second-Order NODEs)的新型跨模态少样本学习方法。通过一个简单而有效的架构——二阶 NODE 模型与跨模态分类器结合,有效缓解了因训练样本少导致的过拟合问题。二阶方法能逼近更广的函数类,提升模型表达力和特征泛化能力。我们使用与类别相关的提示词生成文本嵌入初始化分类器,避免频繁调用文本编码器,提升训练效率;同时采用基于文本的图像增强策略,利用 CLIP 强大的图文关联能力大幅扩充训练数据。在多个数据集上的大量实验表明,SONO 在少样本学习性能上优于现有最先进方法。
原文摘要 · Abstract (English)
We introduce SONO, a novel method leveraging Second-Order Neural Ordinary Differential Equations (Second-Order NODEs) to enhance cross-modal few-shot learning. By employing a simple yet effective architecture consisting of a Second-Order NODEs model paired with a cross-modal classifier, SONO addresses the significant challenge of overfitting, which is common in few-shot scenarios due to limited training examples. Our second-order approach can approximate a broader class of functions, enhancing the model's expressive power and feature generalization capabilities. We initialize our cross-modal classifier with text embeddings derived from class-relevant prompts, streamlining training efficiency by avoiding the need for frequent text encoder processing. Additionally, we utilize text-based image augmentation, exploiting CLIP's robust image-text correlation to enrich training data significantly. Extensive experiments across multiple datasets demonstrate that SONO outperforms existing state-of-the-art methods in few-shot learning performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。