arXiv:2409.13592cs.CVcs.AI2024-09EMNLP被引 12

构建幽默图像理解数据集,测试视觉语言模型的讽刺识别能力

YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models

论文配图:YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models
图 1 · 摘自论文原文
  • 设计三类讽刺理解任务:识别、解释和补全
  • 发布含1084张讽刺图像的高质量数据集YesBut
  • 现有模型在零样本下表现不佳,适合研究幽默理解

理解讽刺与幽默对当前视觉语言模型仍是挑战。本文提出三项新任务:讽刺图像检测(判断图像是否讽刺)、理解(生成图像讽刺的原因)和补全(给定图像一半,从两个选项中选出使整体具有讽刺意味的另一半)。为此,我们发布了高质量数据集YesBut,包含2547张图像,其中1084张为讽刺图像,1463张为非讽刺图像,涵盖多种艺术风格。每张讽刺图像均呈现一个正常场景与一个矛盾场景,构成幽默或反讽效果。尽管当前视觉语言模型在视觉问答、图像描述等多模态任务上表现良好,但在YesBut数据集上的零样本设置下,其性能在自动评估与人工评估中均表现不佳。此外,我们还发布了119张真实的讽刺照片数据集以支持后续研究。数据集与代码已公开于https://github.com/abhi1nandy2/yesbut_dataset。

原文摘要 · Abstract (English)

Understanding satire and humor is a challenging task for even current Vision-Language models. In this paper, we propose the challenging tasks of Satirical Image Detection (detecting whether an image is satirical), Understanding (generating the reason behind the image being satirical), and Completion (given one half of the image, selecting the other half from 2 given options, such that the complete image is satirical) and release a high-quality dataset YesBut, consisting of 2547 images, 1084 satirical and 1463 non-satirical, containing different artistic styles, to evaluate those tasks. Each satirical image in the dataset depicts a normal scenario, along with a conflicting scenario which is funny or ironic. Despite the success of current Vision-Language Models on multimodal tasks such as Visual QA and Image Captioning, our benchmarking experiments show that such models perform poorly on the proposed tasks on the YesBut Dataset in Zero-Shot Settings w.r.t both automated as well as human evaluation. Additionally, we release a dataset of 119 real, satirical photographs for further research. The dataset and code are available at https://github.com/abhi1nandy2/yesbut_dataset.

讽刺理解多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。