无需图文配对数据,就能实现视觉与语言嵌入的无监督匹配。
It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel Data
- 将无监督匹配建模为二次分配问题,提出新启发式算法提升匹配效率。
- 在四个数据集上验证:多数情况下无需标注即可实现有效匹配。
- 证明可构建无监督分类器,零标注下仍具非平凡分类准确率。
柏拉图表示假说认为,随着模型和数据规模增大,视觉与语言嵌入会趋于同质化,各模态内成对距离趋近一致。这表明基础模型成熟后,或可实现完全无监督的跨模态匹配。本文首次开展可行性研究,探究现有视觉与语言基础模型在无监督匹配中的表现。首先,将无监督匹配形式化为二次分配问题,提出新型启发式方法,优于此前求解器;并开发技术以识别高匹配可能性的问题实例。其次,在四个数据集上部署多种视觉与语言模型进行广泛实验。分析显示,许多场景下无需监督即可实现有效匹配。该发现为实现无标注的语义知识迁移提供了可能。作为概念验证,本文展示了一个无监督分类器,在无图像-文本标注条件下仍达到非平凡分类精度。
原文摘要 · Abstract (English)
The platonic representation hypothesis suggests that vision and language embeddings become more homogeneous as model and dataset sizes increase. In particular, pairwise distances within each modality become more similar. This suggests that as foundation models mature, it may become possible to match vision and language embeddings in a fully unsupervised fashion, i.e. without parallel data. We present the first feasibility study, and investigate conformity of existing vision and language foundation models in the context of unsupervised, or "blind", matching. First, we formulate unsupervised matching as a quadratic assignment problem and introduce a novel heuristic that outperforms previous solvers. We also develop a technique to find optimal matching problems, for which a non-trivial match is very likely. Second, we conduct an extensive study deploying a range of vision and language models on four datasets. Our analysis reveals that for many problem instances, vision and language representations can be indeed matched without supervision. This finding opens up the exciting possibility of embedding semantic knowledge into other modalities virtually annotation-free. As a proof of concept, we showcase an unsupervised classifier, which achieves non-trivial classification accuracy without any image-text annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。