arXiv:2502.07391cs.CL2025-02NAACL被引 6

聚焦讽刺目标,提升多模态讽刺解释生成质量

Target-Augmented Shared Fusion-based Multimodal Sarcasm Explanation Generation

  • 引入目标增强的共享融合机制,捕捉图文间讽刺关联
  • 在MORE+数据集上平均提升3.3%,优于现有模型
  • 适合研究讽刺识别与多模态解释生成的学者使用

讽刺是一种旨在以隐含方式嘲讽特定目标(如人物、事件或实体)的语言现象。多模态讽刺解释(MuSE)任务旨在通过自然语言解释揭示讽刺语句中的反讽意图。尽管重要,现有系统忽略了讽刺目标在生成解释中的关键作用。本文提出目标增强的共享融合式讽刺解释模型TURBO,设计新型共享融合机制,利用图像与其标题之间的跨模态关系。TURBO假设讽刺目标,并引导多模态共享融合机制学习反讽背后的细微差别以生成解释。我们在MORE+数据集上评估TURBO模型,相较于多个基线与先进模型,性能平均提升3.3%。此外,我们探索了大语言模型在零样本和单样本设置下的表现,发现其生成的解释虽出色,但常忽略讽刺的核心细节。我们还进行了大规模人工评估,结果表明TURBO生成的解释相较其他系统更优。

原文摘要 · Abstract (English)

Sarcasm is a linguistic phenomenon that intends to ridicule a target (e.g., entity, event, or person) in an inherent way. Multimodal Sarcasm Explanation (MuSE) aims at revealing the intended irony in a sarcastic post using a natural language explanation. Though important, existing systems overlooked the significance of the target of sarcasm in generating explanations. In this paper, we propose a Target-aUgmented shaRed fusion-Based sarcasm explanatiOn model, aka. TURBO. We design a novel shared-fusion mechanism to leverage the inter-modality relationships between an image and its caption. TURBO assumes the target of the sarcasm and guides the multimodal shared fusion mechanism in learning intricacies of the intended irony for explanations. We evaluate our proposed TURBO model on the MORE+ dataset. Comparison against multiple baselines and state-of-the-art models signifies the performance improvement of TURBO by an average margin of $+3.3\%$. Moreover, we explore LLMs in zero and one-shot settings for our task and observe that LLM-generated explanation, though remarkable, often fails to capture the critical nuances of the sarcasm. Furthermore, we supplement our study with extensive human evaluation on TURBO's generated explanations and find them out to be comparatively better than other systems.

多模态讽刺识别生成模型图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。