arXiv:2504.17902cs.CVcs.CL2025-04AAAI被引 1

提升图文仇恨内容识别准确率,关键在精准聚焦有害文本信息

TRACE: Textual Relevance Augmentation and Contextual Encoding for Multimodal Hate Detection

  • 通过视觉上下文增强与标题评分网络,聚焦仇恨相关文本
  • 在Hateful Memes数据集上达0.807准确率和0.806 F1,优于大模型
  • 对多类别仇恨模因泛化能力强,适合真实场景部署

社交媒体表情包是仇恨检测的难点,因其融合视觉与文本线索,表达具文化语境。本文提出TRACE框架,采用分层多模态设计,结合视觉引导的上下文增强、新型标题评分网络以突出仇恨相关文本,并对CLIP文本编码器进行参数高效微调。实验表明,仅微调深层文本编码层显著优于传统投影层微调。该框架在广泛使用的Hateful Memes数据集上达到0.807准确率和0.806 F1-score,性能媲美更大模型且保持高效。同时在MultiOFF Offensive Meme数据集上实现0.673 F1-score,展现出跨类别鲁棒性。分析证实,强视觉对齐与细腻文本表征可有效降低良性干扰导致的误判。代码已公开。

原文摘要 · Abstract (English)

Social media memes are a challenging domain for hate detection because they intertwine visual and textual cues into culturally nuanced messages. To tackle these challenges, we introduce TRACE, a hierarchical multimodal framework that leverages visually grounded context augmentation, along with a novel caption-scoring network to emphasize hate-relevant content, and parameter-efficient fine-tuning of CLIP's text encoder. Our experiments demonstrate that selectively fine-tuning deeper text encoder layers significantly enhances performance compared to simpler projection-layer fine-tuning methods. Specifically, our framework achieves state-of-the-art accuracy (0.807) and F1-score (0.806) on the widely-used Hateful Memes dataset, matching the performance of considerably larger models while maintaining efficiency. Moreover, it achieves superior generalization on the MultiOFF offensive meme dataset (F1-score 0.673), highlighting robustness across meme categories. Additional analyses confirm that robust visual grounding and nuanced text representations significantly reduce errors caused by benign confounders. We publicly release our code to facilitate future research.

多模态检测仇恨内容识别CLIP微调图文融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。