arXiv:2608.26143cs.CLcs.AI2026-08

分析四大视觉语言模型在识别讽刺性仇恨梗图时的推理逻辑与局限。

Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes

论文配图:Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes
图 1 · 摘自论文原文
  • 通过零样本和少样本提示,考察上下文对模型判断的影响。
  • 发现模型常忽略讽刺、语境等细微线索,导致误判。
  • 适合关注AI伦理、内容安全与多模态理解的研究者参考。

梗图已成为人们表达当代社会与政治议题观点的重要工具。其匿名性和病毒式传播特性使其成为仇恨言论的有力载体,但识别这类高度依赖语境的仇恨内容仍极富挑战。尽管视觉语言模型(VLMs)在多模态任务中表现优异,却往往忽视上下文、反讽等关键线索。本文对四种前沿VLMs——LLaVA-7B、Qwen-VL、GPT-4o mini与Claude 3 Haiku——进行定性分析,采用零样本与少样本提示,探究上下文框架如何影响其输出。研究超越传统准确率评估,聚焦模型生成的解释文本,深入揭示其在处理仇恨梗图时的思维过程与内在局限。

原文摘要 · Abstract (English)

Memes have turned out to be a powerful tool through which individuals share their ideas concerning contemporary social and political problems. Their anonymity, as well as their ability to go viral, make them a powerful medium for spreading hate. It remains very difficult to identify such complex and context-dependent hate speech. Although they display excellent performance on multimodal tasks, vision-language models (VLMs) tend to ignore context, irony, and other subtle cues that play a key role in identifying hateful memes. In this work, we present a qualitative analysis of four state-of-the-art VLMs: LLaVA-7B, Qwen-VL, GPT-4o mini, and Claude 3 Haiku. We evaluate these models under zero-shot and few-shot prompting to examine how contextual framing influences their outputs. Our analysis goes beyond simple classification accuracy and focuses on a qualitative evaluation of the models' generated justifications, providing a more in-depth understanding of their thought processes and constraints when dealing with hateful memes.

仇恨言论检测视觉语言模型多模态理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。