arXiv:2607.22016cs.CVcs.AI2026-07被引 1

用多思维链提升图片讽刺内容识别准确率

EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection

论文配图:EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection
图 1 · 摘自论文原文
  • 引入多思维链机制,增强文本与图像的联合推理
  • 在仇恨梗图和多源反讽数据集上准确率达87.3%和91.2%
  • 适合关注跨模态理解与虚假信息检测的研究者

网络梗图常含强烈讽刺或反语,理解其隐含含义需结合图文联合分析。现有方法依赖双流视觉-语言模型提取图文特征,但缺乏背景知识与完整解释能力。简单链式思维(CoT)方法缺乏多视角思考,且依赖浅层特征融合,难以实现细粒度视觉-文本对齐。为此,本文提出增强型视觉-语言多链式思维(EVL-MCoT)框架,通过多思维链提升决策一致性并降低偏差。设计原型引导与上下文引导解码机制,利用视觉原型指导融合过程,实现更精准的图文对齐。在HatefulMemes与MultiOff数据集上取得优异表现,准确率分别达87.3%和91.2%。代码已开源。

原文摘要 · Abstract (English)

MEMEs are widely used on the internet and often carry strong elements of sarcasm or irony. Understanding their hidden meanings typically requires a joint interpretation of text and vision. Existing methods focus on the dual-stream vision-language model to extract the visual and text simultaneously, which lacks background information and prior knowledge about the comprehensive explanation of MEME. One feasible option is to adopt chain-of-thought (CoT). However, the simple CoT approach lacks multi-perspective thinking, which may compromise the reliability of the resulting answers. Moreover, it often relies on shallow feature fusion, lacking the fusion of local details and fine-grained visual-prompt text alignment. This limitation prevents a deeper understanding of the intricate connections between the visual and the text. Herein, an enhanced vision-language multi-CoT (EVL-MCoT) approach is proposed to address these limitations. By promoting multi-CoT, EVL-MCoT enhances consistency and reduces bias in the decision-making process. Additionally, we design a prototype-guided and context-guided decoding framework, which incorporates visual prototypes to guide the fusion process and enables the model to align textual and visual information more precisely. We achieve promising results on the HatefulMemes and MultiOff datasets. The source code has been publicly released and is available at https://github.com/BGWH123/EVL-MCoT.

多模态讽刺识别链式思维图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。