arXiv:2509.21787cs.CVcs.CL2025-09被引 1

用稳定扩散技术识别并模糊图像中的仇恨内容

DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images

  • 结合水印增强的稳定扩散与注意力分析模块定位仇恨区域
  • 生成仇恨注意力图并模糊处理,实现精准去害化
  • 适合社交媒体内容安全与伦理AI研究者使用

网络有害内容激增不仅扭曲公共讨论,也给维护健康数字环境带来挑战。为此,我们构建了一个用于识别数字内容中仇恨信息的多模态数据集。方法核心是创新性地应用带有水印、稳定性增强的稳定扩散技术,并结合数字注意力分析模块(DAAM),精准定位图像中的仇恨元素,生成详细的仇恨注意力图,进而对这些区域进行模糊处理以移除仇恨内容。该数据集作为DeHate共享任务的一部分发布。本文还详述了共享任务的设计。此外,我们提出了DeHater——一个专为多模态去害化任务设计的视觉语言模型。本方法在基于文本提示的图像仇恨检测方面树立新标准,推动社交媒体中更负责任的AI应用发展。

原文摘要 · Abstract (English)

The rise in harmful online content not only distorts public discourse but also poses significant challenges to maintaining a healthy digital environment. In response to this, we introduce a multimodal dataset uniquely crafted for identifying hate in digital content. Central to our methodology is the innovative application of watermarked, stability-enhanced, stable diffusion techniques combined with the Digital Attention Analysis Module (DAAM). This combination is instrumental in pinpointing the hateful elements within images, thereby generating detailed hate attention maps, which are used to blur these regions from the image, thereby removing the hateful sections of the image. We release this data set as a part of the dehate shared task. This paper also describes the details of the shared task. Furthermore, we present DeHater, a vision-language model designed for multimodal dehatification tasks. Our approach sets a new standard in AI-driven image hate detection given textual prompts, contributing to the development of more ethical AI applications in social media.

图像去害扩散模型多模态内容安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。