提出新框架CMAL,用关联学习提升图文预训练效果。
CMAL: A Novel Cross-Modal Associative Learning Framework for Vision-Language Pre-Training
- 通过锚点检测与特征替换构建跨模态关联提示
- 在4个任务上达到先进性能,所需语料更少
- 适合关注高效图文对齐的研究者
随着社交媒体平台的蓬勃发展,视觉-语言预训练(VLP)近年来受到广泛关注,并取得显著进展。其成功主要得益于多模态间的信息互补与增强。然而,当前多数研究聚焦于跨模态对比学习(CMCL),通过拉近正样本对嵌入、推开负样本对嵌入来提升图像-文本对齐,忽略了模态间的自然不对称性,且依赖大规模图文语料才能实现突破。为缓解此问题,我们提出CMAL——一种基于锚点检测与跨模态关联学习的视觉-语言预训练新框架。首先,将视觉物体与文本标记分别嵌入独立超球空间以学习模内隐藏特征;随后设计跨模态关联提示层,通过锚点掩码与特征交换构建混合式跨模态关联提示;接着利用统一语义编码器学习其跨模态交互特征以实现上下文自适应;最后设计关联映射分类层,在锚点处学习模态间的潜在关联映射,并引入新的自监督关联映射分类任务以提升性能。实验结果验证了CMAL的有效性:在四个主流下游视觉-语言任务上表现优于以往基于CMCL的方法,且所需语料显著减少。尤其在SNLI-VE和REC(testA)上取得新最优结果。
原文摘要 · Abstract (English)
With the flourishing of social media platforms, vision-language pre-training (VLP) recently has received great attention and many remarkable progresses have been achieved. The success of VLP largely benefits from the information complementation and enhancement between different modalities. However, most of recent studies focus on cross-modal contrastive learning (CMCL) to promote image-text alignment by pulling embeddings of positive sample pairs together while pushing those of negative pairs apart, which ignores the natural asymmetry property between different modalities and requires large-scale image-text corpus to achieve arduous progress. To mitigate this predicament, we propose CMAL, a Cross-Modal Associative Learning framework with anchor points detection and cross-modal associative learning for VLP. Specifically, we first respectively embed visual objects and textual tokens into separate hypersphere spaces to learn intra-modal hidden features, and then design a cross-modal associative prompt layer to perform anchor point masking and swap feature filling for constructing a hybrid cross-modal associative prompt. Afterwards, we exploit a unified semantic encoder to learn their cross-modal interactive features for context adaptation. Finally, we design an associative mapping classification layer to learn potential associative mappings between modalities at anchor points, within which we develop a fresh self-supervised associative mapping classification task to boost CMAL's performance. Experimental results verify the effectiveness of CMAL, showing that it achieves competitive performance against previous CMCL-based methods on four common downstream vision-and-language tasks, with significantly fewer corpus. Especially, CMAL obtains new state-of-the-art results on SNLI-VE and REC (testA).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。