用多模态融合方法提升数据筛选质量,让少而精的数据训练出更好模型。
Quality over Quantity: Boosting Data Efficiency Through Ensembled Multimodal Data Curation
- 基于多模态操作符的集成框架,自动评分并去重数据点。
- 在38个数据集上平均得分0.182,比基线提升28%。
- 适合追求高效训练的AI研发团队,尤其关注数据质量者。
在数据爆炸的时代,高效清理网络爬取数据集对优化模型性能至关重要。本文针对此类数据结构混乱、异质性强的问题,提出一种学习驱动的先进方法——基于多模态操作符的集成数据筛选框架(EcoDatum)。该方法引入质量引导的去重机制,确保特征分布均衡,并在弱监督集成框架中整合多种单模态与多模态筛选操作符,通过自动化优化为每个数据点打分。EcoDatum显著提升数据筛选质量与效率,在DataComp排行榜上位列第一,38个不同评估数据集上的平均性能得分为0.182,相较DataComp基线方法提升28%,验证了其在数据集构建与模型训练效率方面的有效性。
原文摘要 · Abstract (English)
In an era overwhelmed by vast amounts of data, the effective curation of web-crawl datasets is essential for optimizing model performance. This paper tackles the challenges associated with the unstructured and heterogeneous nature of such datasets. Traditional heuristic curation methods often inadequately capture complex features, resulting in biases and the exclusion of relevant data. We introduce an advanced, learning-driven approach, Ensemble Curation Of DAta ThroUgh Multimodal Operators (EcoDatum), incorporating a novel quality-guided deduplication method to ensure balanced feature distributions. EcoDatum strategically integrates various unimodal and multimodal data curation operators within a weak supervision ensemble framework, utilizing automated optimization to score each data point effectively. EcoDatum, which significantly improves the data curation quality and efficiency, outperforms existing state-of-the-art (SOTA) techniques, ranked 1st on the DataComp leaderboard, with an average performance score of 0.182 across 38 diverse evaluation datasets. This represents a 28% improvement over the DataComp baseline method, demonstrating its effectiveness in improving dataset curation and model training efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。