arXiv:2503.12149cs.CLcs.MM2025-03被引 6

对比12个大模型对多模态讽刺的理解差异,发现模型视角不一且标注有主观性。

Seeing Sarcasm Through Different Eyes: Analyzing Multimodal Sarcasm Perception in Large Vision-Language Models

  • 设计系统化提示词,评估不同大模型对讽刺的多角度理解
  • 在2409个样本上发现模型间和同一模型内判断差异显著
  • 建议用多视角、带不确定性的模型替代单一标签体系

随着大型视觉语言模型(LVLMs)日益具备类人能力,一个关键问题浮现:不同模型是否以不同方式理解多模态讽刺?单个模型能否像人类一样从多个视角把握讽刺?为此,我们基于现有多模态讽刺数据集,设计系统性提示词,构建分析框架。在2409个样本上评估12个顶尖LVLMs,考察模型内部及跨模型的解释差异,重点关注置信度、与数据标签的一致性以及对模糊“中性”案例的识别。进一步在包含多个数据集的100样本小基准上验证,引入扩展提示变体与代表性商业模型。结果揭示显著差异——跨模型及同一模型在不同提示下表现迥异。分类导向提示提升内部一致性,但需进行解释推理时模型分歧明显。这些发现挑战了二元标注范式,凸显讽刺的主观性。我们主张摆脱僵化标注,转向多视角、不确定性感知的建模方式,以深化对多模态讽刺理解的认识。代码与数据已公开于:https://github.com/CoderChen01/LVLMSarcasmAnalysis

原文摘要 · Abstract (English)

With the advent of large vision-language models (LVLMs) demonstrating increasingly human-like abilities, a pivotal question emerges: do different LVLMs interpret multimodal sarcasm differently, and can a single model grasp sarcasm from multiple perspectives like humans? To explore this, we introduce an analytical framework using systematically designed prompts on existing multimodal sarcasm datasets. Evaluating 12 state-of-the-art LVLMs over 2,409 samples, we examine interpretive variations within and across models, focusing on confidence levels, alignment with dataset labels, and recognition of ambiguous "neutral" cases. We further validate our findings on a diverse 100-sample mini-benchmark, incorporating multiple datasets, expanded prompt variants, and representative commercial LVLMs. Our findings reveal notable discrepancies -- across LVLMs and within the same model under varied prompts. While classification-oriented prompts yield higher internal consistency, models diverge markedly when tasked with interpretive reasoning. These results challenge binary labeling paradigms by highlighting sarcasm's subjectivity. We advocate moving beyond rigid annotation schemes toward multi-perspective, uncertainty-aware modeling, offering deeper insights into multimodal sarcasm comprehension. Our code and data are available at: https://github.com/CoderChen01/LVLMSarcasmAnalysis

多模态讽刺识别大模型主观性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。