构建首个视觉隐喻数据集,测试模型理解间接图像线索的能力
A Computational Approach to Visual Metonymy
- 基于符号学理论,用大语言模型与文生图模型生成视觉隐喻图像
- 人类正确率达86.9%,顶尖视觉语言模型仅65.9%,差距明显
- 适合研究多模态推理、认知计算的学者使用
图像常传递超出其字面内容的信息:一组工具可暗示职业,一件文化物品可代表传统。这种间接视觉指代称为视觉隐喻,要求观者通过关联线索推断目标概念,而非直接呈现。本文首次开展视觉隐喻的计算研究,提出一种基于符号学理论的新方法,利用大语言模型与文生图模型生成视觉隐喻图像。基于此框架,我们构建了首个视觉隐喻数据集ViMET,包含2,000道多选题,用于评估多模态语言模型的认知推理能力。在该数据集上的实验表明,人类表现(86.9%)显著优于当前最优视觉-语言模型(65.9%),凸显机器在理解间接视觉指代方面的局限性。数据集已公开:https://github.com/cincynlp/ViMET。
原文摘要 · Abstract (English)
Images often communicate more than they literally depict: a set of tools can suggest an occupation and a cultural artifact can suggest a tradition. This kind of indirect visual reference, known as visual metonymy, invites viewers to recover a target concept via associated cues rather than explicit depiction. In this work, we present the first computational investigation of visual metonymy. We introduce a novel pipeline grounded in semiotic theory that leverages large language models and text-to-image models to generate metonymic visual representations. Using this framework, we construct ViMET, the first visual metonymy dataset comprising 2,000 multiple-choice questions to evaluate the cognitive reasoning abilities in multimodal language models. Experimental results on our dataset reveal a significant gap between human performance (86.9%) and state-of-the-art vision-language models (65.9%), highlighting limitations in machines' ability to interpret indirect visual references. Our dataset is publicly available at: https://github.com/cincynlp/ViMET.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。