arXiv:2506.14473cs.CVcs.LG2025-06ICML被引 2

用多个大模型提升细粒度图像数据选子集效果,比传统方法更优。

Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset Selection

  • 融合多个大模型的互补能力进行数据筛选
  • 在细粒度数据集上超越现有最佳方法
  • 适合需要高效训练的小样本场景

一次性子集选择通过信息提取器(IE)从数据中识别出有代表性的子集,以降低深度学习训练成本。传统IE通常在目标数据集上预训练,具有强数据依赖性。基础模型(FMs)提供了一种新思路,可能缓解此问题。本文研究两个关键问题:(1) 基于FM的子集选择能否在多种数据集上超越传统方法?(2) 所有基础模型在子集选择中表现是否一致?大量实验揭示了意外发现:在细粒度数据集上,FM始终优于传统IE;而在标签噪声较多的粗粒度数据集上,优势减弱。基于此,我们提出RAM-APL(伪类别标签平均准确率排序),专为细粒度图像数据设计。该方法利用多个基础模型的优势互补,显著提升选子集性能,在Oxford-IIIT Pet、Food-101和Caltech-UCSD Birds-200-2011等数据集上达到当前最优水平。

原文摘要 · Abstract (English)

One-shot subset selection serves as an effective tool to reduce deep learning training costs by identifying an informative data subset based on the information extracted by an information extractor (IE). Traditional IEs, typically pre-trained on the target dataset, are inherently dataset-dependent. Foundation models (FMs) offer a promising alternative, potentially mitigating this limitation. This work investigates two key questions: (1) Can FM-based subset selection outperform traditional IE-based methods across diverse datasets? (2) Do all FMs perform equally well as IEs for subset selection? Extensive experiments uncovered surprising insights: FMs consistently outperform traditional IEs on fine-grained datasets, whereas their advantage diminishes on coarse-grained datasets with noisy labels. Motivated by these finding, we propose RAM-APL (RAnking Mean-Accuracy of Pseudo-class Labels), a method tailored for fine-grained image datasets. RAM-APL leverages multiple FMs to enhance subset selection by exploiting their complementary strengths. Our approach achieves state-of-the-art performance on fine-grained datasets, including Oxford-IIIT Pet, Food-101, and Caltech-UCSD Birds-200-2011.

细粒度识别大模型数据筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。