arXiv:2411.03888cs.CL2024-11中稿 · NAACL被引 14

首个跨文化多模态仇恨言论数据集,揭示模型偏见与文化差异。

Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models

  • 构建5语言并行漫画数据集,由跨文化标注者标注。
  • 美印标注者间一致性仅67%,文化差异显著影响判断。
  • 大模型零样本检测更贴近美国标注,存在文化偏差。

全球平台的仇恨言论治理面临多模态、多语言及文化认知差异的挑战。为探究当前视觉-语言模型(VLMs)如何应对这些复杂性,我们构建了首个跨文化、多模态、多语言的仇恨言论平行数据集Multi3Hate,包含5种语言(英语、德语、西班牙语、印地语、中文)的300个并行图文表情包样本,由跨文化标注团队标注。分析显示文化背景对多模态仇恨言论标注有显著影响,各国标注者平均两两一致性仅为74%,远低于随机标注组。定性分析表明,美印标注者间一致性最低,仅67%,归因于文化差异。我们在零样本设置下测试了5个大型VLMs,发现这些模型更倾向于美国标注结果,即使内容以其他文化主导语言呈现。代码与数据集已开源。

原文摘要 · Abstract (English)

Warning: this paper contains content that may be offensive or upsetting Hate speech moderation on global platforms poses unique challenges due to the multimodal and multilingual nature of content, along with the varying cultural perceptions. How well do current vision-language models (VLMs) navigate these nuances? To investigate this, we create the first multimodal and multilingual parallel hate speech dataset, annotated by a multicultural set of annotators, called Multi3Hate. It contains 300 parallel meme samples across 5 languages: English, German, Spanish, Hindi, and Mandarin. We demonstrate that cultural background significantly affects multimodal hate speech annotation in our dataset. The average pairwise agreement among countries is just 74%, significantly lower than that of randomly selected annotator groups. Our qualitative analysis indicates that the lowest pairwise label agreement-only 67% between the USA and India-can be attributed to cultural factors. We then conduct experiments with 5 large VLMs in a zero-shot setting, finding that these models align more closely with annotations from the US than with those from other cultures, even when the memes and prompts are presented in the dominant language of the other culture. Code and dataset are available at https://github.com/MinhDucBui/Multi3Hate.

仇恨言论多模态跨文化VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。