arXiv:2412.20541cs.CLcs.CY2024-12被引 6

构建新型图文仇恨言论数据集,提出结构化推理框架提升 meme 识别准确率。

SAFE-MEME: Structured Reasoning Framework for Robust Hate Speech Detection in Memes

  • 设计双模态思维链与分层分类框架,通过问答式推理增强理解
  • 在两个新数据集上分别提升 2% 和 4.7%,接近闭源模型性能
  • 适配器微调优于全量微调,特别适合细粒度仇恨内容检测

Meme 常作为隐晦传播敏感思想的工具,需上下文知识才能正确解读,导致多模态内容审核困难。现有工作或缺乏高质量细粒度仇恨类别数据集,或依赖低质量社交媒体图像。本文构建了两个新的英文 Meme 仇恨言论数据集:MHS 和 MHS-Con,分别涵盖常规与干扰情境下的细粒度仇恨抽象。我们对多个基线模型进行基准测试,并提出 SAFE-MEME(Structured reAsoning FramEwork)及其两种变体:基于 Q&A 式推理的多模态思维链框架(SAFE-MEME-QA)和分层分类框架(SAFE-MEME-H)。SAFE-MEME-QA 在 MHS 上超越最强开源基线 2%,在 MHS-Con 上提升 4.7%,接近 GPT-4o 与 Gemini 2.5 性能。而 SAFE-MEME-H 进一步超越 SAFE-MEME-QA,相比最佳开源模型提升 3%,在 MHS 上表现接近 GPT-4o。实验表明,在常规细粒度场景下,仅微调单层适配器即优于全模型微调;而在干扰场景中,完整微调配合 Q&A 设计更有效。我们还系统分析错误案例,揭示该框架在鲁棒性与局限性方面的关键洞见。

原文摘要 · Abstract (English)

Memes act as cryptic tools for sharing sensitive ideas, often requiring contextual knowledge to interpret them correctly. It makes multimodal meme moderation difficult, as existing work either lacks high-quality datasets for nuanced hate categories or relies on low-quality social media visuals. Here, we curate two novel multimodal hate speech datasets comprising English memes - MHS and MHS-Con, which capture fine-grained hateful abstractions in regular and confounding scenarios, respectively. We benchmark these datasets against several competing baselines. Furthermore, we introduce SAFE-MEME (Structured reAsoning FramEwork) with its two variants: a novel multimodal Chain-of-Thought based framework employing Q&A-style reasoning (SAFE-MEME-QA) and a hierarchical categorization (SAFE-MEME-H) to enable robust hate speech detection in memes. SAFE-MEME-QA outperforms the strongest open-source baseline model, showing an improvement of 2% on MHS and 4.7% on MHS-Con and closely follows the closed-source models, GPT-4o and Gemini 2.5. In contrast, SAFE-MEME-H surpasses SAFE-MEME-QA, marking a 3% improvement over the best open-source baseline, while achieving performance comparable to GPT-4o only on MHS. We show that fine-tuning a single-layer adapter within SAFE-MEME-H outperforms fully fine-tuned models in regular fine-grained hateful meme detection. However, the fully fine-tuning approach with a Q&A setup is more effective for handling confounding cases. We also systematically examine the error cases, offering valuable insights into the robustness and limitations of the proposed structured reasoning framework for analyzing hateful memes.

仇恨检测图文理解结构化推理Meme 分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。