arXiv:2505.16258cs.CLcs.AI2025-05中稿 · COLM被引 6

用多模态一致性分析提升跨模态讽刺识别能力

IRONIC: Coherence-Aware Reasoning Chains for Multi-Modal Sarcasm Detection

  • 基于图像与文本的指代、类比和语用关联构建推理链
  • 零样本下在多个数据集上达到当前最佳性能
  • 适合研究多模态认知推理与讽刺检测的学者

跨多模态输入理解隐喻语言(如讽刺)具有独特挑战,通常需要任务特定微调和大量推理步骤。然而,现有思维链方法未能有效利用人类识别讽刺时的认知过程。我们提出IRONIC,一种基于上下文学习的框架,通过多模态一致性关系分析图像与文本之间的指代、类比和语用联系。实验表明,IRONIC在不同基线上的零样本多模态讽刺检测中均达到最先进性能,证明了将语言学与认知洞见融入多模态推理设计的重要性。代码已公开于 https://github.com/aashish2000/IRONIC。

原文摘要 · Abstract (English)

Interpreting figurative language such as sarcasm across multi-modal inputs presents unique challenges, often requiring task-specific fine-tuning and extensive reasoning steps. However, current Chain-of-Thought approaches do not efficiently leverage the same cognitive processes that enable humans to identify sarcasm. We present IRONIC, an in-context learning framework that leverages Multi-modal Coherence Relations to analyze referential, analogical and pragmatic image-text linkages. Our experiments show that IRONIC achieves state-of-the-art performance on zero-shot Multi-modal Sarcasm Detection across different baselines. This demonstrates the need for incorporating linguistic and cognitive insights into the design of multi-modal reasoning strategies. Our code is available at: https://github.com/aashish2000/IRONIC

多模态讽刺检测推理链零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。