语言模型能从图像中推断出物体的上位概念,即使没有直接训练数据。
Cross-Modal Taxonomic Generalization in (Vision-) Language Models
- 冻结视觉与语言模型,仅学习中间映射关系。
- 在无显式超类证据时仍可准确预测超类,准确率超70%。
- 视觉相似性高时,跨模态泛化能力依然有效,适合多模态研究者。
语言模型(LM)仅通过表面形式学习语义表示,而视觉-语言模型(VLM)则结合了更具体的视觉证据。本文研究在图像输入条件下,语言模型能否恢复并泛化超类(hypernym)知识。实验采用冻结图像编码器和语言模型的设置,仅训练中间映射层。逐步移除对超类的显式提示,测试模型是否仍能推断。结果表明,在最极端情况下(训练中未提供任何超类证据),模型仍能实现70%以上的准确率。进一步实验显示,当图像-标签配对为反事实但类别内视觉相似度高时,泛化能力仍保持。这说明跨模态分类泛化依赖于外部输入的一致性与语言线索的共同作用。
原文摘要 · Abstract (English)
What is the interplay between semantic representations learned by language models (LM) from surface form alone to those learned from more grounded evidence? We study this question for a scenario where part of the input comes from a different modality -- in our case, in a vision-language model (VLM), where a pretrained LM is aligned with a pretrained image encoder. As a case study, we focus on the task of predicting hypernyms of objects represented in images. We do so in a VLM setup where the image encoder and LM are kept frozen, and only the intermediate mappings are learned. We progressively deprive the VLM of explicit evidence for hypernyms, and test whether knowledge of hypernyms is recoverable from the LM. We find that the LMs we study can recover this knowledge and generalize even in the most extreme version of this experiment (when the model receives no evidence of a hypernym during training). Additional experiments suggest that this cross-modal taxonomic generalization persists under counterfactual image-label mappings only when the counterfactual data have high visual similarity within each category. Taken together, these findings suggest that cross-modal generalization in LMs arises as a result of both coherence in the extralinguistic input and knowledge derived from language cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。