arXiv:2410.01023cs.CVcs.AI2024-10EMNLP被引 11

用视觉笑话测试模型能否结合图像理解文字歧义。

Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!

  • 设计新基准UNPIE,用1000个带图解释的双关语评估多模态理解能力。
  • 图像上下文使模型在复杂任务上的歧义消解准确率显著提升。
  • 适合研究多模态推理、视觉语言模型与语义消歧的研究者。

人类具备多模态素养,能主动融合多种模态信息进行推理。面对文本中的词汇歧义,我们常借助缩略图或教科书插图等其他模态补充理解。机器能否实现类似能力?为此,我们提出理解图像解释的双关语(UNPIE)基准,用于评估多模态输入在解决词汇歧义中的作用。双关语因其固有的歧义性成为理想测试对象。数据集包含1000个双关语,每个均配有解释两种含义的图像。我们设置三个多模态挑战:双关语定位(Pun Grounding)、歧义消解(Disambiguation)和重构(Reconstruction),并提供标注以评估多模态素养的不同方面。结果表明,在引入视觉上下文后,各类Socratic模型和视觉语言模型的表现优于纯文本模型,尤其在任务复杂度提高时优势更明显。

原文摘要 · Abstract (English)

Humans possess multimodal literacy, allowing them to actively integrate information from various modalities to form reasoning. Faced with challenges like lexical ambiguity in text, we supplement this with other modalities, such as thumbnail images or textbook illustrations. Is it possible for machines to achieve a similar multimodal understanding capability? In response, we present Understanding Pun with Image Explanations (UNPIE), a novel benchmark designed to assess the impact of multimodal inputs in resolving lexical ambiguities. Puns serve as the ideal subject for this evaluation due to their intrinsic ambiguity. Our dataset includes 1,000 puns, each accompanied by an image that explains both meanings. We pose three multimodal challenges with the annotations to assess different aspects of multimodal literacy; Pun Grounding, Disambiguation, and Reconstruction. The results indicate that various Socratic Models and Visual-Language Models improve over the text-only models when given visual context, particularly as the complexity of the tasks increases.

多模态双关语视觉语言模型歧义消解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。