TADS通过任务感知筛选数据,用更少数据实现更好多模态模型性能。
TADS: Task-Aware Data Selection for Multi-Task Multimodal Pre-Training
- 融合质量、任务相关性与多样性,构建可学习的数据选择函数。
- 仅用36%数据即超越基线平均1.0%,零样本性能显著提升。
- 适合追求数据高效训练的多任务多模态研究者使用。
大型多模态预训练模型如CLIP依赖高质量训练数据,但原始网络爬取数据常含噪声、对齐错误和冗余,导致训练效率低且泛化能力差。现有数据筛选方法或基于启发式规则,存在偏差且多样性不足;或为数据驱动但任务无关,无法优化多任务场景。为此,我们提出TADS(Task-Aware Data Selection),一种面向多任务多模态预训练的新框架,将内在质量、任务相关性和分布多样性整合进可学习的价值函数。TADS采用包含单模态与跨模态操作的综合质量评估系统,通过可解释的相似性向量量化任务相关性,并以基于聚类的加权策略优化多样性。通过反馈驱动的元学习机制,根据多个下游任务的代理模型表现自适应优化筛选策略。在CC12M数据集上的实验表明,TADS仅使用36%的数据,就在ImageNet、CIFAR-100、MS-COCO和Flickr30K等基准上实现更优的零样本性能,平均超越基线1.0%。这表明TADS通过筛选高价值数据子集,在相同计算约束下显著提升性能上限。
原文摘要 · Abstract (English)
Large-scale multimodal pre-trained models like CLIP rely heavily on high-quality training data, yet raw web-crawled datasets are often noisy, misaligned, and redundant, leading to inefficient training and suboptimal generalization. Existing data selection methods are either heuristic-based, suffering from bias and limited diversity, or data-driven but task-agnostic, failing to optimize for multi-task scenarios. To address these gaps, we introduce TADS (Task-Aware Data Selection), a novel framework for multi-task multimodal pre-training that integrates Intrinsic Quality, Task Relevance, and Distributional Diversity into a learnable value function. TADS employs a comprehensive quality assessment system with unimodal and cross-modal operators, quantifies task relevance via interpretable similarity vectors, and optimizes diversity through cluster-based weighting. A feedback-driven meta-learning mechanism adaptively refines the selection strategy based on proxy model performance across multiple downstream tasks. Experiments on CC12M demonstrate that TADS achieves superior zero-shot performance on benchmarks like ImageNet, CIFAR-100, MS-COCO, and Flickr30K, using only 36% of the data while outperforming baselines by an average of 1.0%. This highlights that TADS significantly enhances data efficiency by curating a high-utility subset that yields a much higher performance ceiling within the same computational constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。