首个面向非洲语言的公平安全评估基准,解决文化错配与低资源语言安全短板。
UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages
- 基于155位专家设计的对抗性查询构建文化适配的安全策略
- 多语言模型在非洲语言上表现不佳,动态策略仍难完全本地化
- 适合关注AI公平性、多语言安全与非洲语种研究者参考
当前守护模型主要基于西方视角并针对高资源语言优化,使低资源非洲语言易受新型危害、跨语言失效及文化偏差影响。多数模型依赖僵化的预设安全类别,难以适应多样语言与社会文化背景。为实现鲁棒安全,需具备灵活、可运行时执行的政策与反映本地规范、风险场景和文化期待的评估基准。我们提出UbuntuGuard,首个基于政策的非洲语言安全评估基准,由来自医疗等敏感领域共155位专家撰写的对抗性查询构建。从中提炼出情境化安全策略与参考响应,捕捉文化相关的风险信号,支持政策对齐的模型评估。我们评估了15个模型(7个通用大模型与8个守护模型),涵盖静态、动态与多语言三种变体。结果表明:现有以英语为中心的基准高估了多语言实际安全水平;跨语言迁移仅提供部分覆盖;动态模型虽能在推理时更好利用策略,但仍难以充分本地化非洲语言上下文。研究凸显了发展多语言、文化适配的安全基准的紧迫性,以推动低资源语言可靠且公平的守护模型发展。
原文摘要 · Abstract (English)
Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerable to evolving harms, cross-lingual failures, and cultural misalignment. Moreover, most guardian models rely on rigid, predefined safety categories that fail to generalize across diverse linguistic and sociocultural contexts. Achieving robust safety requires flexible, runtime-enforceable policies and benchmarks that reflect local norms, harm scenarios, and cultural expectations. We introduce UbuntuGuard, the first policy-based safety benchmark for African languages built from adversarial queries authored by 155 domain experts across sensitive fields, including healthcare. From these expert-crafted queries, we derive context-specific safety policies and reference responses that capture culturally grounded risk signals, enabling policy-aligned evaluation of guardian models. We evaluate 15 models, comprising seven general-purpose LLMs and eight guardian models across three distinct variants: static, dynamic, and multilingual. Our findings reveal that existing English-centric benchmarks overestimate real-world multilingual safety, cross-lingual transfer provides partial but insufficient coverage, and dynamic models, while better equipped to leverage policies at inference time, still struggle to fully localize African-language contexts. These findings highlight the urgent need for multilingual, culturally grounded safety benchmarks to enable the development of reliable and equitable guardian models for low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。