arXiv:2601.04692cs.CLcs.CV2026-01被引 3

用生成式AI实现仇恨梗图的检测、解释与提前干预

See, Explain, and Intervene: A Few-Shot Multimodal Agent Framework for Hateful Meme Moderation

  • 构建多模态代理框架,利用生成模型实现少样本适配
  • 在有限标注数据下仍能有效识别并解释仇恨梗图
  • 适合需要快速部署的平台内容安全系统

本文从检测、解释和提前干预三个角度研究仇恨梗图问题,提出基于生成式AI的多模态代理框架。鉴于大规模标注数据成本高昂,该框架利用任务特定的生成式多模态代理与大模型的少样本适应能力,应对不同类型的梗图。这是首个聚焦于低数据条件下可泛化仇恨梗图治理的工作,具备实际生产部署潜力。警告:内容可能包含有害信息。

原文摘要 · Abstract (English)

In this work, we examine hateful memes from three complementary angles - how to detect them, how to explain their content and how to intervene them prior to being posted - by applying a range of strategies built on top of generative AI models. To the best of our knowledge, explanation and intervention have typically been studied separately from detection, which does not reflect real-world conditions. Further, since curating large annotated datasets for meme moderation is prohibitively expensive, we propose a novel framework that leverages task-specific generative multimodal agents and the few-shot adaptability of large multimodal models to cater to different types of memes. We believe this is the first work focused on generalizable hateful meme moderation under limited data conditions, and has strong potential for deployment in real-world production scenarios. Warning: Contains potentially toxic contents.

仇恨内容多模态生成模型少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。