arXiv:2507.19551cs.CYcs.AI2025-07被引 9

测试了针对LGBTQ+内容的有害梗图检测模型鲁棒性,发现文本扰动是主要弱点。

Rainbow Noise: Stress-Testing Multimodal Harmful-Meme Detectors on LGBTQ Content

  • 构建首个针对LGBTQ+有害梗图的鲁棒性评测基准,涵盖图文扰动组合
  • 现有模型对文本修改敏感,尤其MemeBLIP2在语义篡改下性能下降超40%
  • 轻量级文本去噪适配器(TDA)显著提升模型抗干扰能力,适合安全系统优化

针对针对LGBTQ+群体的仇恨梗图常通过修改标题、图像或两者来逃避检测。本文构建首个该场景下的鲁棒性评测基准,将四种真实标题攻击与三种典型图像退化相结合,在PrideMM数据集上测试所有组合。以MemeCLIP和MemeBLIP2两个先进检测器为案例研究,提出轻量级文本去噪适配器(TDA)增强后者鲁棒性。实验表明,尽管所有系统均依赖文本信息,但架构设计和预训练数据显著影响其稳定性;其中MemeCLIP表现更稳健,而MemeBLIP2对文本篡改极度敏感。引入TDA后,其整体鲁棒性超越其他模型。消融分析显示,当前多模态安全模型在面对精准文本扰动时仍存在明显脆弱性,而如TDA等轻量级模块可成为有效防御路径。

原文摘要 · Abstract (English)

Hateful memes aimed at LGBTQ\,+ communities often evade detection by tweaking either the caption, the image, or both. We build the first robustness benchmark for this setting, pairing four realistic caption attacks with three canonical image corruptions and testing all combinations on the PrideMM dataset. Two state-of-the-art detectors, MemeCLIP and MemeBLIP2, serve as case studies, and we introduce a lightweight \textbf{Text Denoising Adapter (TDA)} to enhance the latter's resilience. Across the grid, MemeCLIP degrades more gently, while MemeBLIP2 is particularly sensitive to the caption edits that disrupt its language processing. However, the addition of the TDA not only remedies this weakness but makes MemeBLIP2 the most robust model overall. Ablations reveal that all systems lean heavily on text, but architectural choices and pre-training data significantly impact robustness. Our benchmark exposes where current multimodal safety models crack and demonstrates that targeted, lightweight modules like the TDA offer a powerful path towards stronger defences.

多模态安全有害内容检测鲁棒性评测LGBTQ+

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。