用文字描述降低视觉预测不确定性,让模型学得更懂语义的视觉表征。
Text-Conditional JEPA for Learning Semantically Rich Visual Representations

- 用文本条件调节图像块特征预测,减少遮蔽区域的模糊性。
- 在多个任务上超越对比学习方法,尤其擅长细粒度视觉理解。
- 无需对比损失,仅靠特征预测实现强大预训练,适合视觉语言任务。
基于图像的联合嵌入预测架构(I-JEPA)通过掩码特征预测提供了一种有前景的视觉自监督学习方法。然而,由于掩码位置存在固有的视觉不确定性,特征预测仍具挑战,可能无法学习到语义表征。本文提出文本条件JEPA(TC-JEPA),利用图像字幕减少预测不确定性。具体而言,我们使用细粒度文本条件器对预测的图像块特征进行调制,该条件器在输入文本标记上计算稀疏交叉注意力。通过这种条件化,图像块特征可被表示为文本的函数,从而更具语义意义。实验表明,TC-JEPA提升了下游任务性能和训练稳定性,并展现出良好的扩展性。此外,该方法还提出了仅基于特征预测的新视觉-语言预训练范式,在多样任务中表现优于对比方法,尤其是在需要细粒度视觉理解与推理的任务上。
原文摘要 · Abstract (English)
Image-based Joint-Embedding Predictive Architecture (I-JEPA) offers a promising approach to visual self-supervised learning through masked feature prediction. However with the inherent visual uncertainty at masked positions, feature prediction remains challenging and may fail to learn semantic representations. In this work, we propose Text-Conditional JEPA (TC-JEPA) that uses image captions to reduce the prediction uncertainty. Specifically, we modulate the predicted patch features using a fine-grained text conditioner that computes sparse cross-attention over input text tokens. With such conditioning, patch features become predictable as a function of text, thus are more semantically meaningful. We show TC-JEPA improves downstream performance and training stability, with promising scaling properties. TC-JEPA also offers a new vision-language pretraining paradigm based on feature prediction only, outperforming contrastive methods on diverse tasks, especially those requiring fine-grained visual understanding and reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。