综述自监督对比学习在图文分析中的进展与应用
A Survey on Self-supervised Contrastive Learning for Multimodal Text-Image Analysis
- 通过构建图文正负样本对,无须人工标注即可学习语义特征
- 涵盖多种模型结构与最新预训练任务设计方法
- 适合关注多模态自监督学习的研究者与工程师
自监督学习通过从无标签数据中挖掘内在模式来生成隐式标签,并提取判别性特征。对比学习引入正负样本概念:正样本对(如同一图像的不同变换)在嵌入空间中靠近,负样本对(如不同图像/对象的视图)则被拉远。该方法显著提升了图像理解与图文分析能力,且对标注数据依赖极低。本文系统梳理了对比学习在图文模型中的术语、近年发展与应用场景。首先概述近年来图文模型中对比学习的主流方法;其次按模型结构进行分类;再次深入讨论预训练任务设计、网络架构及关键趋势;最后总结自监督对比学习在图文模型中的前沿应用。
原文摘要 · Abstract (English)
Self-supervised learning is a machine learning approach that generates implicit labels by learning underlined patterns and extracting discriminative features from unlabeled data without manual labelling. Contrastive learning introduces the concept of "positive" and "negative" samples, where positive pairs (e.g., variation of the same image/object) are brought together in the embedding space, and negative pairs (e.g., views from different images/objects) are pushed farther away. This methodology has shown significant improvements in image understanding and image text analysis without much reliance on labeled data. In this paper, we comprehensively discuss the terminologies, recent developments and applications of contrastive learning with respect to text-image models. Specifically, we provide an overview of the approaches of contrastive learning in text-image models in recent years. Secondly, we categorize the approaches based on different model structures. Thirdly, we further introduce and discuss the latest advances of the techniques used in the process such as pretext tasks for both images and text, architectural structures, and key trends. Lastly, we discuss the recent state-of-art applications of self-supervised contrastive learning Text-Image based models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。