arXiv:2508.03654cs.CLcs.CV2025-08中稿 · CIKM 2025被引 10

评测大模型对多模态讽刺的理解能力,提出无需训练的改进框架

Can Large Vision-Language Models Understand Multimodal Sarcasm?

  • 引入物体提取与外部概念知识增强视觉理解
  • 多模型实验验证框架显著提升讽刺识别与解释能力
  • 适合关注多模态情感分析、大模型可解释性的研究者

讽刺是一种复杂的语言现象,其字面意义与实际意图存在差异,给情感分析等情绪敏感任务带来挑战。传统讽刺检测方法主要依赖文本,近年研究开始融合多模态信息。然而,大型视觉语言模型(LVLMs)在多模态讽刺分析(MSA)中的应用仍不充分。本文评估了LVLMs在多模态讽刺检测与解释任务中的表现,发现其存在视觉理解不足和概念知识缺乏等关键局限。为此,我们提出一种无需训练的框架,通过深度物体提取与外部概念知识融合,提升模型在多模态情境下对讽刺的解读与解释能力。在多个模型上的实验结果表明,该框架有效。代码已公开于https://github.com/cp-cp/LVLM-MSA。

原文摘要 · Abstract (English)

Sarcasm is a complex linguistic phenomenon that involves a disparity between literal and intended meanings, making it challenging for sentiment analysis and other emotion-sensitive tasks. While traditional sarcasm detection methods primarily focus on text, recent approaches have incorporated multimodal information. However, the application of Large Visual Language Models (LVLMs) in Multimodal Sarcasm Analysis (MSA) remains underexplored. In this paper, we evaluate LVLMs in MSA tasks, specifically focusing on Multimodal Sarcasm Detection and Multimodal Sarcasm Explanation. Through comprehensive experiments, we identify key limitations, such as insufficient visual understanding and a lack of conceptual knowledge. To address these issues, we propose a training-free framework that integrates in-depth object extraction and external conceptual knowledge to improve the model's ability to interpret and explain sarcasm in multimodal contexts. The experimental results on multiple models show the effectiveness of our proposed framework. The code is available at https://github.com/cp-cp/LVLM-MSA.

多模态讽刺检测视觉语言模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。