用改进的图文数据提升对比学习,让模型理解图像中物体间的组合关系。
Learning Visual Composition through Improved Semantic Guidance
- 通过优化弱标注图文数据,增强标准对比学习
- 在组合性任务上性能显著超越原有CLIP模型和专用架构
- 适合研究视觉理解、多模态学习的开发者参考
视觉图像并非孤立对象的集合,而是多种流动概念的组合体现。尽管视觉表征学习取得显著进展,但这些方法多聚焦于少数离散物体的表征,缺乏对物体间交互关系的理解。现有基于描述文本或对比学习的模型往往将图像视为‘词汇袋’,忽略组合结构。一些工作尝试设计专用架构来解决这一问题,但复杂且难扩展。本文提出简单可扩展的方法:通过大幅改进弱标注数据(即图像描述)质量,显著提升标准对比学习的表现。先前的CLIP模型在测试组合性能力的任务中表现接近随机水平,而本文方法大幅提升其性能,超过所有专用架构。我们在基于DOCCI的新描述基准上验证结果,通过一系列消融实验表明,仅用增强数据训练的标准CLIP模型,在图像检索任务中已具备出色表现。
原文摘要 · Abstract (English)
Visual imagery does not consist of solitary objects, but instead reflects the composition of a multitude of fluid concepts. While there have been great advances in visual representation learning, such advances have focused on building better representations for a small number of discrete objects bereft of an understanding of how these objects are interacting. One can observe this limitation in representations learned through captions or contrastive learning -- where the learned model treats an image essentially as a bag of words. Several works have attempted to address this limitation through the development of bespoke learned architectures to directly address the shortcomings in compositional learning. In this work, we focus on simple, and scalable approaches. In particular, we demonstrate that by substantially improving weakly labeled data, i.e. captions, we can vastly improve the performance of standard contrastive learning approaches. Previous CLIP models achieved near chance rate on challenging tasks probing compositional learning. However, our simple approach boosts performance of CLIP substantially and surpasses all bespoke architectures. Furthermore, we showcase our results on a relatively new captioning benchmark derived from DOCCI. We demonstrate through a series of ablations that a standard CLIP model trained with enhanced data may demonstrate impressive performance on image retrieval tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。