用大模型生成可解释分析,提升仇恨梗图识别准确率与可信度
Demystifying Hateful Content: Leveraging Large Multimodal Models for Hateful Meme Detection with Explainable Decisions
- 利用大模型生成人类式解读,结合图文编码提升判断力
- 在三个数据集上超越现有最佳模型,显著减少误判
- 适合需要透明决策的平台内容审核场景
仇恨梗图检测作为多模态任务,因隐含仇恨信息和上下文线索复杂而极具挑战。以往方法依赖预训练视觉语言模型(PT-VLMs)的隐含知识与注意力机制,但其决策过程不透明,难以解释,影响可信度。本文提出IntMeme框架,利用大模型(LMMs)生成人类式可解释分析,深入理解图文内容与语境。该框架采用独立编码模块分别处理梗图及其解释,并融合特征以提升分类性能。实验在三个数据集上验证了IntMeme的有效性,显著优于当前最优模型,缓解了原有模型的黑箱问题与误判现象。
原文摘要 · Abstract (English)
Hateful meme detection presents a significant challenge as a multimodal task due to the complexity of interpreting implicit hate messages and contextual cues within memes. Previous approaches have fine-tuned pre-trained vision-language models (PT-VLMs), leveraging the knowledge they gained during pre-training and their attention mechanisms to understand meme content. However, the reliance of these models on implicit knowledge and complex attention mechanisms renders their decisions difficult to explain, which is crucial for building trust in meme classification. In this paper, we introduce IntMeme, a novel framework that leverages Large Multimodal Models (LMMs) for hateful meme classification with explainable decisions. IntMeme addresses the dual challenges of improving both accuracy and explainability in meme moderation. The framework uses LMMs to generate human-like, interpretive analyses of memes, providing deeper insights into multimodal content and context. Additionally, it uses independent encoding modules for both memes and their interpretations, which are then combined to enhance classification performance. Our approach addresses the opacity and misclassification issues associated with PT-VLMs, optimizing the use of LMMs for hateful meme detection. We demonstrate the effectiveness of IntMeme through comprehensive experiments across three datasets, showcasing its superiority over state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。