CLIP通过预训练中物体出现频率预测组合泛化能力,揭示其可重组组件提升现实任务表现。
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
- 基于物体在预训练中的出现频率,预测组合新样本的性能
- 频率低的物体组合仍能准确预测,说明模型具备组件重组能力
- 平衡物体频次可提升泛化性,无需增加数据量
我们通过性能预测研究了CLIP模型在真实世界数据上实现组合泛化的成功条件。已有研究表明,CLIP在单个概念上的线性性能提升需要指数级增长的预训练数据,这种样本效率低下可能通过将新输入系统性理解为已学组件的组合得以缓解,使罕见观察映射到常见概念。为探索CLIP的组合泛化能力,我们筛选出预训练语料库中未出现的物体组合样本。结果表明,这些样本上CLIP的性能可由个体物体在预训练中的出现频率准确预测。研究发现,CLIP能从预训练数据中解耦物体,并能直接重组它们。此外,我们首次展示了该能力随预训练数据规模的扩展规律。实际数据筛选中,平衡物体出现频次可提升泛化性能,有助于在不扩大数据量的前提下提高效率与准确性。
原文摘要 · Abstract (English)
We investigate the success conditions for compositional generalization of CLIP models on real-world data through performance prediction. Prior work shows that CLIP requires exponentially more pretraining data for linear performance gains on individual concepts. This sample-inefficient scaling could be mitigated if CLIP systematically understood new inputs as compositions of learned components, allowing rare observation to be mapped to common concepts. To explore CLIP's compositional generalization ability, we filter retrieval corpora for samples with object combinations not present in the pretraining corpus. We show that CLIP's performance on these samples can be accurately predicted from the pretraining frequencies of individual objects. Our findings demonstrate that CLIP learns to disentangle objects observed in its pretraining data and can recompose them straightforwardly. Additionally, we are the first to show how this ability scales with pretraining data. For data curation in practice, our results suggest that balancing object occurrences improves generalization, which should benefit CLIP's efficiency and accuracy without scaling data volume.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。