通过识别误判风险模式,提升对隐性有害表情包的检测准确率。
Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval
- 基于误判风险模式构建知识库,动态引导大模型推理
- 在5个任务上平均提升8.3%的F1分数和7.7%准确率
- 特别适合处理对抗性与未见过的隐性有害内容
网络表情包已成为流行的多模态媒介,但正被日益用于通过反讽、隐喻等微妙修辞传递有害观点。现有检测方法,包括基于多模态大语言模型(MLLM)的技术,难以应对这些隐含表达,导致频繁误判。本文提出PatMD,一种通过学习并主动规避潜在误判风险来检测有害表情包的新方法。核心思路是超越表面内容匹配,识别深层误判风险模式,主动引导MLLM避开已知误判陷阱。我们首先构建知识库,将每个表情包拆解为解释其可能被误判的原因(漏检或误判)的风险模式。针对目标表情包,PatMD检索相关模式并动态引导MLLM推理。在包含6,626张表情包的基准数据集上,5项有害检测任务实验表明,PatMD优于当前最优基线,平均提升8.30% F1-score与7.71%准确率,且在未见及对抗性样本上表现一致稳健。
原文摘要 · Abstract (English)
Internet memes have emerged as a popular multimodal medium, yet they are increasingly weaponized to convey harmful opinions through subtle rhetorical devices like irony and metaphor. Existing detection approaches, including Multimodal Large Language Model (MLLM)-based techniques, struggle with these implicit expressions, leading to frequent misjudgments. This paper introduces PatMD, a novel approach that detects harmful memes by learning from and proactively mitigating these potential misjudgment risks. Our core idea is to move beyond superficial content-level matching and instead identify the underlying misjudgment risk patterns, proactively guiding the MLLMs to avoid known misjudgment pitfalls. We first construct a knowledge base where each meme is deconstructed into a misjudgment risk pattern explaining why it might be misjudged, either overlooking harmful undertones (false negative) or overinterpreting benign content (false positive). For a given target meme, PatMD retrieves relevant patterns and utilizes them to dynamically guide the MLLM's reasoning. Experiments on a benchmark of 6,626 memes across 5 harmful detection tasks show that PatMD outperforms state-of-the-art baselines, achieving an average of 8.30% improvement in F1-score and 7.71% improvement in accuracy, while exhibiting consistent robustness on unseen and adversarial memes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。