arXiv:2501.08506cs.LGcs.AI2025-01

研究发现数据多样性比模型大小更能提升视觉模型性能。

Exploring the Efficacy of Meta-Learning: Unveiling Superior Data Diversity Utilization of MAML Over Pre-training

  • 用任务向量量化数据多样性,分析其对模型影响
  • 在12个数据集上验证准确率与多样性呈中强正相关(R²=0.15-0.42)
  • 适合关注数据质量而非单纯堆数据的研究者

当前大模型训练主要关注数据与模型规模,但对数据集其他属性的影响缺乏探索。本文假设数据多样性会影响视觉模型性能,并在12个主流视觉数据集(如Omniglot、CIFAR-FS、Aircraft)和5种模型配置下,对比预训练与模型无关元学习方法。结果表明,测试准确率与数据多样性存在中等到强的正相关(R²: 0.15–0.42),损失与多样性亦有弱但显著的相关性(R²≈0.2)。研究支持了数据多样性的重要性,证明使用(Task2Vec)度量数据多样性是理解大规模学习的关键,强调了深入理解数据对构建更强大、泛化能力更强模型的价值。

原文摘要 · Abstract (English)

Currently, data and model size dominate the narrative in the training of super-large, powerful models. However, there has been a lack of exploration on the effect of other attributes of the training dataset on model performance. We hypothesize that dataset diversity can impact the performance of vision models. Our study shows positive correlations between test set accuracy and data diversity, providing an argument for furthering the research of dataset attributes beyond size. We analyzed pre-training and model-agnostic meta-learning methods on twelve popular visual datasets (e.g., Omniglot, CIFAR-FS, Aircraft) and five model configurations, including MAML variants with different numbers of inner gradient steps and supervised learning. We show moderate to strong positive correlations (R-squared: 0.15-0.42) between accuracy and data diversity and weaker but significant correlations (R-squared: ~0.2) between loss and diversity. These findings support our hypothesis and demonstrate a promising way for a deeper exploration of how formal data diversity influences model performance. This initial study highlights the potential of (Task2Vec) data diversity as a valuable measure in the rapidly evolving field of large-scale learning and emphasizes that understanding the dataset is key to building more powerful and generalizable models.

元学习数据多样性模型性能视觉任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。