构建多模态跨语言幽默危害检测基准,区分显性与隐性伤害。
Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
- 设计含3000文本、6000图像、1200视频的多语言数据集
- 闭源模型在英阿语上表现显著优于开源模型
- 聚焦隐性伤害识别,推动文化敏感的安全对齐
黑色幽默常依赖微妙的文化细节和隐含线索,需上下文推理才能理解,现有静态基准难以捕捉此类安全挑战。为此,我们提出首个多模态、跨语言的有害幽默检测与理解基准。数据集包含3000条文本、6000张图像(英语与阿拉伯语),以及1200个视频(涵盖英语、阿拉伯语及无语言依赖的通用场景)。不同于常规毒性数据集,我们采用严格标注规范:将笑话分为安全类、有害类;有害类进一步细分为显性(直接)与隐性(间接)两类,以探测深层推理能力。我们在所有模态上系统评估了当前最先进的开闭源模型。结果表明,闭源模型显著优于开源模型,且英阿语间性能差异明显,凸显文化背景与推理能力在安全对齐中的关键作用。提醒:本文包含可能冒犯、有害或偏见的内容。
原文摘要 · Abstract (English)
Dark humor often relies on subtle cultural nuances and implicit cues that require contextual reasoning to interpret, posing safety challenges that current static benchmarks fail to capture. To address this, we introduce a novel multimodal, multilingual benchmark for detecting and understanding harmful and offensive humor. Our manually curated dataset comprises 3,000 texts and 6,000 images in English and Arabic, alongside 1,200 videos that span English, Arabic, and language-independent (universal) contexts. Unlike standard toxicity datasets, we enforce a strict annotation guideline: distinguishing Safe jokes from Harmful ones, with the latter further classified into Explicit (overt) and Implicit (Covert) categories to probe deep reasoning. We systematically evaluate state-of-the-art (SOTA) open and closed-source models across all modalities. Our findings reveal that closed-source models significantly outperform open-source ones, with a notable difference in performance between the English and Arabic languages in both, underscoring the critical need for culturally grounded, reasoning-aware safety alignment. Warning: this paper contains example data that may be offensive, harmful, or biased.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。