arXiv:2505.24190cs.LGcs.CV2025-05ICML被引 4

用理论指导生成合成数据,提升少样本模型泛化能力

Provably Improving Generalization of Few-Shot Models with Synthetic Data

  • 基于分布差距理论设计合成数据生成与训练策略
  • 在多个数据集上超越现有最优方法,显著提升少样本分类性能
  • 适合研究少样本学习与数据增强的学者与工程师

少样本图像分类因标注样本稀缺而面临挑战。用合成数据扩充训练样本成为缓解此问题的有前景方法,但由真实与合成数据分布差异导致的性能下降仍是个难题。本文提出一个理论框架,量化分布偏差对监督学习的影响,尤其针对图像分类任务。该框架进一步指导如何生成优质合成样本并训练具备高泛化能力的预测器。基于此,我们提出一种新型理论驱动算法,结合原型学习优化数据划分与模型训练,有效弥合真实少样本数据与合成数据间的差距。大量实验结果表明,该方法在多个数据集上均优于当前最优方法,展现出卓越性能。

原文摘要 · Abstract (English)

Few-shot image classification remains challenging due to the scarcity of labeled training examples. Augmenting them with synthetic data has emerged as a promising way to alleviate this issue, but models trained on synthetic samples often face performance degradation due to the inherent gap between real and synthetic distributions. To address this limitation, we develop a theoretical framework that quantifies the impact of such distribution discrepancies on supervised learning, specifically in the context of image classification. More importantly, our framework suggests practical ways to generate good synthetic samples and to train a predictor with high generalization ability. Building upon this framework, we propose a novel theoretical-based algorithm that integrates prototype learning to optimize both data partitioning and model training, effectively bridging the gap between real few-shot data and synthetic data. Extensive experiments results show that our approach demonstrates superior performance compared to state-of-the-art methods, outperforming them across multiple datasets.

少样本学习合成数据泛化能力原型学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。