arXiv:2409.05136cs.CL2024-09被引 13

提出多模态仇恨言论检测新框架,提升识别准确率。

MHS-STMA: Multimodal Hate Speech Detection via Scalable Transformer-Based Multilevel Attention Framework

  • 用多层注意力机制融合文本与图像信息
  • 在三个数据集上均超越基线模型表现
  • 适合关注社交媒体内容安全的研究者

社交媒体对人们生活影响深远,仇恨言论已成为严重社会问题。文本与图像作为主要的多模态数据形式,在文章中广泛分布。早期方法多聚焦单模态分析,而多模态研究常忽视各模态的独特性。为此,本文提出一种可扩展的多模态仇恨内容检测架构——基于变压器的多层级注意力(STMA)。该框架包含三部分:联合注意力深度学习机制、视觉注意力编码器和标题注意力编码器。各组件通过不同注意力过程处理多模态数据,分别捕捉跨模态关联。在Hateful Meme、MultiOff和MMHS150K三个仇恨言论数据集上,采用多种评估标准验证了该方法的有效性。结果表明,所提方法在所有数据集上均优于基准模型。

原文摘要 · Abstract (English)

Social media has a significant impact on people's lives. Hate speech on social media has emerged as one of society's most serious issues in recent years. Text and pictures are two forms of multimodal data that are distributed within articles. Unimodal analysis has been the primary emphasis of earlier approaches. Additionally, when doing multimodal analysis, researchers neglect to preserve the distinctive qualities associated with each modality. To address these shortcomings, the present article suggests a scalable architecture for multimodal hate content detection called transformer-based multilevel attention (STMA). This architecture consists of three main parts: a combined attention-based deep learning mechanism, a vision attention-mechanism encoder, and a caption attention-mechanism encoder. To identify hate content, each component uses various attention processes and handles multimodal data in a unique way. Several studies employing multiple assessment criteria on three hate speech datasets such as Hateful memes, MultiOff, and MMHS150K, validate the suggested architecture's efficacy. The outcomes demonstrate that on all three datasets, the suggested strategy performs better than the baseline approaches.

多模态仇恨言论注意力机制文本图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。