用视觉语言模型识别并转化仇恨型图文梗图,保障网络环境安全。
Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models
- 基于定义提示的检测方法,提升仇恨内容识别准确率。
- 提出UnHateMeme框架,可替换图文中的仇恨元素。
- 效果达人类标准,适合内容安全与AI伦理研究者使用。
社交媒体快速发展为用户创作内容提供了新渠道,多模态梗图常用于幽默表达,但有时被滥用于传播针对个人或群体的仇恨言论。尽管仇恨梗图检测已有较多研究,但有效转化其内容仍具挑战。本文利用视觉语言模型(VLMs)的强大生成与推理能力,提出两项关键贡献:一是基于定义的提示技术用于检测仇恨梗图;二是名为UnHateMeme的统一框架,通过替换文本和/或视觉中的仇恨成分实现内容缓解。该方法在主流预训练VLMs如LLaVA、Gemini和GPT-4o上均表现出色,能将仇恨梗图转化为符合人类判断标准且图文一致的非仇恨形式。实验验证了其有效性,为构建更安全的在线环境提供了重要路径。
原文摘要 · Abstract (English)
The rapid evolution of social media has provided enhanced communication channels for individuals to create online content, enabling them to express their thoughts and opinions. Multimodal memes, often utilized for playful or humorous expressions with visual and textual elements, are sometimes misused to disseminate hate speech against individuals or groups. While the detection of hateful memes is well-researched, developing effective methods to transform hateful content in memes remains a significant challenge. Leveraging the powerful generation and reasoning capabilities of Vision-Language Models (VLMs), we address the tasks of detecting and mitigating hateful content. This paper presents two key contributions: first, a definition-guided prompting technique for detecting hateful memes, and second, a unified framework for mitigating hateful content in memes, named UnHateMeme, which works by replacing hateful textual and/or visual components. With our definition-guided prompts, VLMs achieve impressive performance on hateful memes detection task. Furthermore, our UnHateMeme framework, integrated with VLMs, demonstrates a strong capability to convert hateful memes into non-hateful forms that meet human-level criteria for hate speech and maintain multimodal coherence between image and text. Through empirical experiments, we show the effectiveness of state-of-the-art pretrained VLMs such as LLaVA, Gemini and GPT-4o on the proposed tasks, providing a comprehensive analysis of their respective strengths and limitations for these tasks. This paper aims to shed light on important applications of VLMs for ensuring safe and respectful online environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。