用知识图谱高效搜集数据,用1000万张图训练出高质量生物领域CLIP模型
Using Knowledge Graphs to harvest datasets for efficient CLIP model training
- 借助知识图谱优化网页搜索,智能筛选训练数据
- 仅用1000万图像即可训练出针对生物领域的专家级CLIP模型
- 构建3300万图像+4600万文本的EntityNet数据集,加速通用模型训练
训练高质量CLIP模型通常需要海量数据,这限制了特定领域模型的发展——尤其在现有最大规模CLIP模型覆盖不足的领域——并显著增加训练成本。这对需要精细控制训练过程的科学研究构成挑战。本文展示,通过结合知识图谱的智能网络搜索策略,可仅用较少数据从头训练出稳健的CLIP模型。具体而言,我们仅用1000万张图像就构建了一个面向生物体的专家基础模型。此外,我们推出了EntityNet数据集,包含3300万张图像与4600万条文本描述,使通用CLIP模型的训练时间大幅缩短。
原文摘要 · Abstract (English)
Training high-quality CLIP models typically requires enormous datasets, which limits the development of domain-specific models -- especially in areas that even the largest CLIP models do not cover well -- and drives up training costs. This poses challenges for scientific research that needs fine-grained control over the training procedure of CLIP models. In this work, we show that by employing smart web search strategies enhanced with knowledge graphs, a robust CLIP model can be trained from scratch with considerably less data. Specifically, we demonstrate that an expert foundation model for living organisms can be built using just 10M images. Moreover, we introduce EntityNet, a dataset comprising 33M images paired with 46M text descriptions, which enables the training of a generic CLIP model in significantly reduced time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。