arXiv:2502.13061cs.CLcs.AI2025-02EMNLP被引 19

提升大模型对仇恨梗图的鲁棒检测能力,兼顾准确率与跨域泛化。

Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection

  • 引入稳健适配框架,优化多模态模型在仇恨梗图识别中的表现。
  • 在6个数据集上达到最优性能,优于更大规模的智能系统。
  • 生成更高质量推理链,提升模型可解释性,适合安全审计场景。

仇恨梗图已成为网络重大问题,亟需可靠的自动化检测系统。尽管大视觉语言模型(LMMs)在该任务中展现出潜力,但仍存在性能不佳和跨领域泛化能力弱的问题。近期研究进一步揭示了监督微调(SFT)与上下文学习在该场景下的局限性。为此,本文提出一种稳健的适配框架,旨在提升模型在域内检测精度与跨域泛化能力的同时,保持其通用视觉-语言理解能力。分析表明,该方法在对抗攻击下比SFT模型更具鲁棒性。在六个梗图分类数据集上的实验显示,本方法性能达当前最优,超越更大规模的代理系统。此外,该方法生成的推理链条质量更高,有助于解释仇恨内容,增强模型可解释性。代码已开源:https://github.com/JingbiaoMei/RGCL。

原文摘要 · Abstract (English)

Hateful memes have become a significant concern on the Internet, necessitating robust automated detection systems. While Large Multimodal Models (LMMs) have shown promise in hateful meme detection, they face notable challenges like sub-optimal performance and limited out-of-domain generalization capabilities. Recent studies further reveal the limitations of both supervised fine-tuning (SFT) and in-context learning when applied to LMMs in this setting. To address these issues, we propose a robust adaptation framework for hateful meme detection that enhances in-domain accuracy and cross-domain generalization while preserving the general vision-language capabilities of LMMs. Analysis reveals that our approach achieves improved robustness under adversarial attacks compared to SFT models. Experiments on six meme classification datasets show that our approach achieves state-of-the-art performance, outperforming larger agentic systems. Moreover, our method generates higher-quality rationales for explaining hateful content compared to standard SFT, enhancing model interpretability. Code available at https://github.com/JingbiaoMei/RGCL

仇恨内容检测多模态模型鲁棒性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。