arXiv:2602.19212cs.CL2026-02

针对低资源孟加拉语仇恨表情包,提出融合检索增强的双注意力模型。

Retrieval Augmented Enhanced Dual Co-Attention Framework for Target Aware Multimodal Bengali Hateful Meme Detection

  • 用检索增强数据提升样本多样性与类别平衡
  • 新模型在仇恨识别和目标检测上分别达0.78和0.71准确率
  • 适合研究低资源语言仇恨内容检测的学者

社交媒体中的仇恨内容日益以图文结合的表情包形式出现。在孟加拉语等低资源语言中,自动检测面临标注数据少、类别不平衡和广泛混码等问题。本文通过引入孟加拉语多模态攻击数据集(MIMOSA)中语义对齐的样本,扩充了孟加拉语仇恨表情包数据集(BHM),提升了类别平衡与语义多样性。提出增强型双协同注意力框架(xDORA),融合视觉编码器(CLIP、DINOv2)与多语言文本编码器(XGLM、XLM-R),通过加权注意力池化学习鲁棒跨模态表征。基于这些嵌入,构建基于FAISS的k近邻分类器实现非参数推理,并引入RAG-Fused DORA,结合检索驱动的上下文推理。进一步在零样本、少样本及检索增强提示下评估LLaVA。实验表明,xDORA(CLIP + XLM-R)在仇恨表情包识别和目标实体检测上分别获得0.78和0.71的宏平均F1分数,而RAG-Fused DORA进一步提升至0.79和0.74,优于基线。FAISS分类器表现良好,对罕见类别具有鲁棒性。相比之下,LLaVA在少样本设置下效果有限,检索增强仅带来小幅提升,凸显预训练视觉-语言模型在未微调的混码孟加拉语内容上的局限性。结果表明,监督式、检索增强及非参数化多模态框架在应对低资源仇恨话语的语言文化复杂性方面有效。

原文摘要 · Abstract (English)

Hateful content on social media increasingly appears as multimodal memes that combine images and text to convey harmful narratives. In low-resource languages such as Bengali, automated detection remains challenging due to limited annotated data, class imbalance, and pervasive code-mixing. To address these issues, we augment the Bengali Hateful Memes (BHM) dataset with semantically aligned samples from the Multimodal Aggression Dataset in Bengali (MIMOSA), improving both class balance and semantic diversity. We propose the Enhanced Dual Co-attention Framework (xDORA), integrating vision encoders (CLIP, DINOv2) and multilingual text encoders (XGLM, XLM-R) via weighted attention pooling to learn robust cross-modal representations. Building on these embeddings, we develop a FAISS-based k-nearest neighbor classifier for non-parametric inference and introduce RAG-Fused DORA, which incorporates retrieval-driven contextual reasoning. We further evaluate LLaVA under zero-shot, few-shot, and retrieval-augmented prompting settings. Experiments on the extended dataset show that xDORA (CLIP + XLM-R) achieves macro-average F1-scores of 0.78 for hateful meme identification and 0.71 for target entity detection, while RAG-Fused DORA improves performance to 0.79 and 0.74, yielding gains over the DORA baseline. The FAISS-based classifier performs competitively and demonstrates robustness for rare classes through semantic similarity modeling. In contrast, LLaVA exhibits limited effectiveness in few-shot settings, with only modest improvements under retrieval augmentation, highlighting constraints of pretrained vision-language models for code-mixed Bengali content without fine-tuning. These findings demonstrate the effectiveness of supervised, retrieval-augmented, and non-parametric multimodal frameworks for addressing linguistic and cultural complexities in low-resource hate speech detection.

仇恨内容检测多模态低资源语言检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。