arXiv:2510.11852cs.LG2025-10中稿 · ICDMW 2025 Worksho…被引 4

评测7个开源视觉语言模型在多模态讽刺检测上的表现

Evaluating Open-Source Vision-Language Models for Multimodal Sarcasm Detection

  • 用零样本、少样本提示评估模型对图文讽刺的识别能力
  • 模型在讽刺检测上表现中等,但解释生成质量不高
  • 适合研究多模态讽刺或可解释AI的读者参考

近期开源视觉语言模型(VLMs)为理解复杂主观的多模态现象(如讽刺)提供了新机遇。本文评估了七种先进VLMs——BLIP2、InstructBLIP、OpenFlamingo、LLaVA、PaliGemma、Gemma3和Qwen-VL——在零样本、单样本和少样本提示下检测多模态讽刺的能力,并评估其生成讽刺实例解释的能力。在三个基准讽刺数据集(Muse、MMSD2.0和SarcNet)上进行测试。主要目标为:(1) 量化各模型对讽刺图文对的检测性能;(2) 评估其生成能揭示视觉-文本矛盾的人类级解释的能力。结果显示,当前模型在二分类讽刺检测上取得中等成效,但未经过任务特定微调时,仍无法生成高质量解释。

原文摘要 · Abstract (English)

Recent advances in open-source vision-language models (VLMs) offer new opportunities for understanding complex and subjective multimodal phenomena such as sarcasm. In this work, we evaluate seven state-of-the-art VLMs - BLIP2, InstructBLIP, OpenFlamingo, LLaVA, PaliGemma, Gemma3, and Qwen-VL - on their ability to detect multimodal sarcasm using zero-, one-, and few-shot prompting. Furthermore, we evaluate the models' capabilities in generating explanations to sarcastic instances. We evaluate the capabilities of VLMs on three benchmark sarcasm datasets (Muse, MMSD2.0, and SarcNet). Our primary objectives are twofold: (1) to quantify each model's performance in detecting sarcastic image-caption pairs, and (2) to assess their ability to generate human-quality explanations that highlight the visual-textual incongruities driving sarcasm. Our results indicate that, while current models achieve moderate success in binary sarcasm detection, they are still not able to generate high-quality explanations without task-specific finetuning.

多模态讽刺检测视觉语言模型解释生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。