无需难例训练,用文本反演提升模型组合理解能力
A New Method to Capturing Compositional Knowledge in Linguistic Space
- 通过文本反演将图像映射为伪标记,实现零样本组合理解
- 在SugarCREPE上超越现有模型8%以上,图像检索也显著提升
- 引入'no'逻辑正则化与知识蒸馏,解决标记交互和计算耗时问题
组合理解使视觉语言模型能够解析图像与文本中对象、属性和关系的复杂关联。然而,现有方法多依赖难例样本与微调,易高估性能且难获取难例。本文提出零样本组合理解(ZS-CU)新任务,不需硬负样本训练。我们提出YUKINO(通过文本反演获得组合理解知识,无需硬负样本),利用文本反演将未标注图像映射至预训练CLIP模型的伪标记空间,并引入'no'逻辑正则化缓解标记交互问题。此外,采用知识蒸馏降低文本反演的时间复杂度。实验表明,YUKINO在SugarCREPE基准上比现有多模态最先进模型提升超8%,图像检索任务亦有显著改善。
原文摘要 · Abstract (English)
Compositional understanding allows visual language models to interpret complex relationships between objects, attributes, and relations in images and text. However, most existing methods often rely on hard negative examples and fine-tuning, which can overestimate improvements and are limited by the difficulty of obtaining hard negatives. In this work, we introduce Zero-Shot Compositional Understanding (ZS-CU), a novel task that enhances compositional understanding without requiring hard negative training data. We propose YUKINO (Yielded Compositional Understanding Knowledge via Textual Inversion with NO), which uses textual inversion to map unlabeled images to pseudo-tokens in a pre-trained CLIP model. We propose introducing "no" logical regularization to address the issue of token interaction in inversion. Additionally, we suggest using knowledge distillation to reduce the time complexity of textual inversion. Experimental results show that YUKINO outperforms the existing multi-modal SOTA models by over 8% on the SugarCREPE benchmark, and also achieves significant improvements in image retrieval tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。