用多模态模型识别并解释网络迷因中的性别歧视内容
What is Beneath Misogyny: Misogynous Memes Classification and Explanation
- 分模态处理图文,用交叉注意力融合特征
- 在新数据集上准确分类四类性别刻板印象
- 可解释原因,适合研究网络偏见的学者
迷因在现代广泛传播,常以娱乐形式出现,但可能暗藏性别歧视等有害意识形态。由于其多模态特性(图像与文字)及在不同社会语境下的隐晦表现,检测和理解迷因中的性别歧视极具挑战。本文提出一种新型多模态方法MM-Misogyny,分别处理文本与图像模态,并通过交叉注意力机制统一为多模态上下文,再由分类器与大语言模型实现标签生成、类别划分与解释。评估基于新构建的数据集WBMS(What's Beneath Misogynous Stereotyping),该数据集从网络收集了刻板化性别歧视迷因,分为厨房、领导力、工作与购物四类。实验表明,该方法在检测与分类上优于现有方法,且能提供对性别歧视在日常生活领域中运作机制的细粒度理解。代码与数据集已开源。
原文摘要 · Abstract (English)
Memes are popular in the modern world and are distributed primarily for entertainment. However, harmful ideologies such as misogyny can be propagated through innocent-looking memes. The detection and understanding of why a meme is misogynous is a research challenge due to its multimodal nature (image and text) and its nuanced manifestations across different societal contexts. We introduce a novel multimodal approach, \textit{namely}, \textit{\textbf{MM-Misogyny}} to detect, categorize, and explain misogynistic content in memes. \textit{\textbf{MM-Misogyny}} processes text and image modalities separately and unifies them into a multimodal context through a cross-attention mechanism. The resulting multimodal context is then easily processed for labeling, categorization, and explanation via a classifier and Large Language Model (LLM). The evaluation of the proposed model is performed on a newly curated dataset (\textit{\textbf{W}hat's \textbf{B}eneath \textbf{M}isogynous \textbf{S}tereotyping (WBMS)}) created by collecting misogynous memes from cyberspace and categorizing them into four categories, \textit{namely}, Kitchen, Leadership, Working, and Shopping. The model not only detects and classifies misogyny, but also provides a granular understanding of how misogyny operates in domains of life. The results demonstrate the superiority of our approach compared to existing methods. The code and dataset are available at \href{https://github.com/kushalkanwarNS/WhatisBeneathMisogyny/tree/main}{https://github.com/Misogyny}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。