测试视觉语言模型对广告中非常规图像的推理能力
Benchmarking VLMs' Reasoning About Persuasive Atypical Images
- 设计三类新任务评估模型对非常规图像的理解
- 发现模型推理能力远弱于语言模型,但可提取关键特征
- 揭示非常规元素有助于理解广告说服意图,适合广告分析研究者
视觉语言模型(VLMs)在多项任务中表现出强大的零样本泛化能力,尤其在与大语言模型(LLMs)结合后。然而,它们对修辞性、说服性视觉媒体(如广告)的理解仍缺乏研究。广告常使用非常规图像,通过意外的物体组合传达共性特征。例如,图1(e)中啤酒呈现羽毛质感,需推理出其象征酒体轻盈。为此,我们提出三项新任务:多标签非常规分类、非常规语句检索和非常规物体识别,用于评测VLMs对说服性图像中非常规性的理解能力。通过引入语义挑战性负例,检验模型的推理表现,并首创基于非常规性的描述生成方法,提取敏感于非常规元素的全面图像描述。结果表明:(1) 与LLMs相比,VLMs缺乏高级推理能力;(2) 简单有效策略可提取非常规感知信息,实现完整图像描述;(3) 非常规性有助于理解广告说服逻辑。代码与数据将公开。
原文摘要 · Abstract (English)
Vision language models (VLMs) have shown strong zero-shot generalization across various tasks, especially when integrated with large language models (LLMs). However, their ability to comprehend rhetorical and persuasive visual media, such as advertisements, remains understudied. Ads often employ atypical imagery, using surprising object juxtapositions to convey shared properties. For example, Fig. 1 (e) shows a beer with a feather-like texture. This requires advanced reasoning to deduce that this atypical representation signifies the beer's lightness. We introduce three novel tasks, Multi-label Atypicality Classification, Atypicality Statement Retrieval, and Aypical Object Recognition, to benchmark VLMs' understanding of atypicality in persuasive images. We evaluate how well VLMs use atypicality to infer an ad's message and test their reasoning abilities by employing semantically challenging negatives. Finally, we pioneer atypicality-aware verbalization by extracting comprehensive image descriptions sensitive to atypical elements. Our findings reveal that: (1) VLMs lack advanced reasoning capabilities compared to LLMs; (2) simple, effective strategies can extract atypicality-aware information, leading to comprehensive image verbalization; (3) atypicality aids persuasive advertisement understanding. Code and data will be made available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。