arXiv:2507.04458cs.CL2025-07被引 4

用双重推理专家提升多模态讽刺识别准确率

Think Twice Before You Judge: Mixture of Dual Reasoning Experts for Multimodal Sarcasm Detection

  • 融合内部矛盾检测与外部思维链推理,动态加权选择最佳路径
  • 在两个基准数据集上超越现有模型,显著提升识别效果
  • 适合需要理解深层语境的社交媒体内容分析任务

多模态讽刺检测因社交媒体多媒体内容兴起而受到关注。理解图像-文本讽刺内容常需外部背景知识,如文化参考或常识推理。然而,现有模型难以捕捉讽刺背后的深层逻辑,主要依赖图像标题或对象属性等浅层线索。为此,我们提出MiDRE(Mixture of Dual Reasoning Experts),集成内部推理专家以检测图像-文本对内的不一致,以及外部推理专家,利用大视觉语言模型通过思维链提示生成的结构化推理过程。自适应门控机制动态权衡两个专家,选择最相关推理路径。不同于将外部知识视为静态输入的方法,MiDRE仅在外部知识有益时启用,降低大模型幻觉或无关信号风险。在两个基准数据集上的实验表明,MiDRE优于基线模型。定性分析显示,即使外部推理偶尔存在噪声,仍能提供关键线索,引导模型更好理解讽刺。

原文摘要 · Abstract (English)

Multimodal sarcasm detection has attracted growing interest due to the rise of multimedia posts on social media. Understanding sarcastic image-text posts often requires external contextual knowledge, such as cultural references or commonsense reasoning. However, existing models struggle to capture the deeper rationale behind sarcasm, relying mainly on shallow cues like image captions or object-attribute pairs from images. To address this, we propose \textbf{MiDRE} (\textbf{Mi}xture of \textbf{D}ual \textbf{R}easoning \textbf{E}xperts), which integrates an internal reasoning expert for detecting incongruities within the image-text pair and an external reasoning expert that utilizes structured rationales generated via Chain-of-Thought prompting to a Large Vision-Language Model. An adaptive gating mechanism dynamically weighs the two experts, selecting the most relevant reasoning path. Unlike prior methods that treat external knowledge as static input, MiDRE selectively adapts to when such knowledge is beneficial, mitigating the risks of hallucinated or irrelevant signals from large models. Experiments on two benchmark datasets show that MiDRE achieves superior performance over baselines. Various qualitative analyses highlight the crucial role of external rationales, revealing that even when they are occasionally noisy, they provide valuable cues that guide the model toward a better understanding of sarcasm.

多模态讽刺检测思维链视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。