让视觉语言模型读懂抽象描述,提升图文检索效果。
Seeing the Abstract: Translating the Abstract Language for Vision Language Models
- 用预训练模型和已有数据,无须训练即可转换抽象语义到具体表示。
- 在跨数据集图文检索中,性能超越微调过的模型,提升显著。
- 适用于各类视觉语言模型,适合希望快速增强模型理解力的研究者。
自然语言不仅描述视觉内容,更包含情感、创意等难以直接感知的抽象概念。当前视觉语言模型(VLMs)对这类抽象语言关注不足。本研究首次系统分析发现,在时尚领域这一高度抽象表达的典型场景中,抽象词汇占比与具体词汇相当,且能提供新信息,对检索任务有实际帮助。然而,现有通用或时尚专用VLMs的训练语料缺乏足够抽象词,导致其无法有效表征抽象语言。为此,提出无需训练、适配所有模型的抽象转具体翻译器(ACT),利用预训练模型和现有多模态数据库,在隐空间中将抽象表示映射至具体表示。实验表明,即使不进行训练,ACT在文本到图像检索任务中仍优于微调后的VLMs,且在不同模型和数据集间均表现一致,具备强泛化能力,可即插即用。
原文摘要 · Abstract (English)
Natural language goes beyond dryly describing visual content. It contains rich abstract concepts to express feeling, creativity and properties that cannot be directly perceived. Yet, current research in Vision Language Models (VLMs) has not shed light on abstract-oriented language. Our research breaks new ground by uncovering its wide presence and under-estimated value, with extensive analysis. Particularly, we focus our investigation on the fashion domain, a highly-representative field with abstract expressions. By analyzing recent large-scale multimodal fashion datasets, we find that abstract terms have a dominant presence, rivaling the concrete ones, providing novel information, and being useful in the retrieval task. However, a critical challenge emerges: current general-purpose or fashion-specific VLMs are pre-trained with databases that lack sufficient abstract words in their text corpora, thus hindering their ability to effectively represent abstract-oriented language. We propose a training-free and model-agnostic method, Abstract-to-Concrete Translator (ACT), to shift abstract representations towards well-represented concrete ones in the VLM latent space, using pre-trained models and existing multimodal databases. On the text-to-image retrieval task, despite being training-free, ACT outperforms the fine-tuned VLMs in both same- and cross-dataset settings, exhibiting its effectiveness with a strong generalization capability. Moreover, the improvement introduced by ACT is consistent with various VLMs, making it a plug-and-play solution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。