arXiv:2508.09175cs.CVcs.AI2025-08被引 25

用多模态图神经网络精准识别针对女性的侮辱内容

A Context-aware Attention and Graph Neural Network-based Multimodal Framework for Misogyny Detection

  • 引入上下文感知注意力与图结构重构特征,增强跨模态理解
  • 在两个数据集上分别提升10.17%和8.88%的宏平均F1值
  • 适合研究性别歧视检测、社交媒体安全的学者与工程师

社交媒体中大量仇恨内容针对女性。现有通用内容检测方法难以有效识别厌女言论,亟需专门解决方案。本文提出一种基于多模态图神经网络的厌女内容检测框架,包含三个模块:多模态注意力模块(MANM)通过自适应门控机制聚焦关键文本与视觉信息;基于图的特征重建模块(GFRM)在单模态内优化特征表示;内容特定特征学习模块(CFLM)捕捉毒性特征与图文关联特征。此外,构建厌女词典以计算文本中的针对性得分,并采用测试时特征空间增强提升泛化能力。在包含11,000样本的MAMI和13,494样本的MMHS150K数据集上验证,相比现有方法,宏观F1值分别提升10.17%和8.88%。

原文摘要 · Abstract (English)

A substantial portion of offensive content on social media is directed towards women. Since the approaches for general offensive content detection face a challenge in detecting misogynistic content, it requires solutions tailored to address offensive content against women. To this end, we propose a novel multimodal framework for the detection of misogynistic and sexist content. The framework comprises three modules: the Multimodal Attention module (MANM), the Graph-based Feature Reconstruction Module (GFRM), and the Content-specific Features Learning Module (CFLM). The MANM employs adaptive gating-based multimodal context-aware attention, enabling the model to focus on relevant visual and textual information and generating contextually relevant features. The GFRM module utilizes graphs to refine features within individual modalities, while the CFLM focuses on learning text and image-specific features such as toxicity features and caption features. Additionally, we curate a set of misogynous lexicons to compute the misogyny-specific lexicon score from the text. We apply test-time augmentation in feature space to better generalize the predictions on diverse inputs. The performance of the proposed approach has been evaluated on two multimodal datasets, MAMI and MMHS150K, with 11,000 and 13,494 samples, respectively. The proposed method demonstrates an average improvement of 10.17% and 8.88% in macro-F1 over existing methods on the MAMI and MMHS150K datasets, respectively.

厌女检测多模态图神经网络社会安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。