arXiv:2602.07497cs.CL2026-02被引 3

测试视觉语言模型在跨文化仇恨梗图检测中的表现

From Native Memes to Global Moderation: Cross-Cultural Evaluation of Vision-Language Models for Hateful Meme Detection

  • 用多语言梗图数据集测试模型在不同文化语境下的表现
  • 本地语言提示和少量样本学习显著提升检测效果
  • 提醒警惕翻译后检测的偏差,适合做内容安全研究者

文化背景深刻影响人们对网络内容的理解,但现有视觉语言模型(VLMs)大多基于西方或英语中心的数据训练,限制了其在仇恨梗图检测等任务中的公平性与跨文化鲁棒性。本文提出系统评估框架,诊断并量化先进VLMs在多语言梗图数据集上的跨文化表现,分析三个维度:(i) 学习策略(零样本与少样本),(ii) 提示语言(本地语言与英文),(iii) 翻译对语义与检测的影响。结果表明,常见的‘翻译后检测’方法会降低性能,而采用本地语言提示和少样本学习则显著提升检测能力。研究揭示了模型系统性趋同于西方安全标准的现象,并提出可操作的缓解策略,为构建全球鲁棒的多模态内容审核系统提供指导。

原文摘要 · Abstract (English)

Cultural context profoundly shapes how people interpret online content, yet vision-language models (VLMs) remain predominantly trained through Western or English-centric lenses. This limits their fairness and cross-cultural robustness in tasks like hateful meme detection. We introduce a systematic evaluation framework designed to diagnose and quantify the cross-cultural robustness of state-of-the-art VLMs across multilingual meme datasets, analyzing three axes: (i) learning strategy (zero-shot vs. one-shot), (ii) prompting language (native vs. English), and (iii) translation effects on meaning and detection. Results show that the common ``translate-then-detect'' approach deteriorate performance, while culturally aligned interventions - native-language prompting and one-shot learning - significantly enhance detection. Our findings reveal systematic convergence toward Western safety norms and provide actionable strategies to mitigate such bias, guiding the design of globally robust multimodal moderation systems.

视觉语言模型跨文化仇恨内容检测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。