首个阿拉伯语仇恨表情包细粒度标注数据集,助力文化语境下仇恨内容识别。
AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes

- 构建5000张阿拉伯语仇恨表情包的细粒度多标签数据集
- 发现文化隐喻与图文协同是识别关键,现有模型表现有限
- 适合研究跨文化仇恨内容检测、社会计算与多模态安全的学者
仇恨表情包是一种日益严重的多模态网络危害,其敌意意图常需结合图像、文本、文化引用及隐含目标共同理解。尽管高资源语言的仇恨表情包检测已有进展,阿拉伯语仍缺乏系统研究,现有资源多聚焦于宣传内容或粗粒度有害标签。我们提出AHA-Memes(阿拉伯语仇恨表情包),据我们所知,这是首个大规模、细粒度、多标签标注的阿拉伯语仇恨表情包基准数据集。该数据集包含5000张人工标注的表情包,采用捕捉攻击策略的分类体系;此外还提供约66000张银标签数据以支持后续研究。我们对仅文本、仅图像、后期融合的多模态模型,以及少样本上下文学习(ICL)和开/闭源视觉-语言模型(VLMs)在零样本与微调设置下进行了基准测试。结果建立了强基线,并揭示了文化语境下阿拉伯语仇恨表情包检测的关键挑战。我们已公开数据集、标注指南与评估脚本,以推动未来研究。警告:本文包含可能令人不适的内容。
原文摘要 · Abstract (English)
Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cultural references, and implicit targets. While hateful meme detection has advanced in high-resource languages, Arabic remains underexplored, with existing meme resources focusing mainly on propaganda or coarse harmful-content labels. We introduce AHA-Memes (Arabic HAteful Memes), which is, to our knowledge, the first large-scale Arabic hateful meme benchmark with fine-grained, multi-label annotations. The dataset includes 5K manually annotated memes using a taxonomy that captures hate types, i.e., attack strategies. We further provide ~66K silver-labeled memes to support future studies. We benchmark text-only, image-only, and late-fusion multimodal models, as well as few-shot in-context learning (ICL) and open- and closed-weight Vision-Language Models (VLMs) under zero-shot and fine-tuning settings. Our results establish strong baselines and highlight key challenges in culturally grounded Arabic hateful meme detection. We release the dataset, annotation guidelines, and evaluation scripts to support future research. WARNING: This paper contains examples that may be disturbing to readers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。