arXiv:2607.06196cs.CLcs.CY2026-07

构建首个跨文化多模态安全评估基准,揭示AI在不同地区的真实风险

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

论文配图:Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
图 1 · 摘自论文原文
  • 从文化出发构建6国8语的多模态安全数据集,真实反映本地禁忌
  • 发现模型在特定地区常因图像与语境组合触发法律/文化违规,但全局指标无法捕捉
  • 提出基于共识的AI裁判系统,助力识别地域化失效模式,适合全球AI安全研究者

当前AI安全评估多依赖西方中心、文化中立的默认标准,忽视区域法律、社会语言差异和文化禁忌,导致视觉语言模型(VLMs)在全球部署中存在漏洞。我们提出Pluralis v0.1:一个从文化优先视角构建的多模态、多区域、多语言数据集,覆盖亚太六国(孟加拉、印度、韩国、巴基斯坦、新加坡、台湾)共6,448个提示,涉及八种语言。不同于以往将西方数据适配至本地的做法,Pluralis原生采集本地化安全风险。其创新在于引入多模态评估范式:用户文本(如“我该送这个吗?”)与指向“这个”的图像(如时钟)单独无害,但组合后可能触发特定法律或文化违规。该设计分离了普适性安全问题与本地文化适宜性,将其列为首要评估维度。为此,我们构建Judge-Pluralis——一个基于共识机制的大型语言模型裁判集合,训练于实证推导的文化分类体系。在子集测试中,观察到VLM存在系统性本地化失败模式,如图像误识致下游伤害、忽略物品-语境-地域交互、拒绝响应不足等,且这些模式随地区和语言系统性变化,而全局平均指标难以揭示。最终,Pluralis并非已完成的评估框架,而是推动未来多语言、多文化评估创新的起点,呼吁研究界共同推进全球AI文化对齐科学。

原文摘要 · Abstract (English)

Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances, and cultural taboos, leaving Vision-Language Models (VLMs) vulnerable in global deployments. We introduce Pluralis v0.1: a novel multimodal, multi-regional, and multilingual dataset built from a culture-first perspective. Spanning 6,448 prompts across six Asia-Pacific countries (Bangladesh, India, Korea, Pakistan, Singapore, Taiwan) and eight languages, Pluralis diverges from prior work by natively sourcing localized safety hazards rather than adapting Western datasets. Crucially, it introduces a multimodal evaluation paradigm: user text (e.g., "Should I gift this?") and an image referring to "this" (e.g., a clock) - both innocuous in isolation, but synergistically triggering specific legal or cultural violations. Pluralis disentangles universal safety violations from localized cultural appropriateness, establishing the latter as a first-class evaluation axis. To operationalize this, we present Judge-Pluralis, an agreement-gated LLM-as-a-Judge ensemble trained on examples classified in an empirically derived cultural taxonomy. Observing VLM behavior on a subset of the Pluralis surfaces recurring, locale-specific failure modes such as image misidentifications with downstream harm, missed item-context-locale interactions, and inadequate refusals. These failure modes vary systematically across locales and languages, exposing blind spots that globally averaged metrics conceal. Ultimately, Pluralis is not presented as a solved evaluation framework for cultural alignment, but rather as a first step and catalyst for future innovation. We call upon the research community to utilize this foundation to advance the science of multilingual, multicultural evaluation to better support AI cultural alignment globally.

多模态评估文化对齐AI安全多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。