首个面向东南亚语言与文化的AI安全评估基准,揭示主流模型在本地语境下表现显著下降。
SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures

- 构建覆盖8种东南亚语言的原生人工验证安全数据集
- 在本地化场景中,顶尖大模型性能比英文测试下降超30%
- 适合关注多语言公平性、区域化AI安全的研究者与开发者
防护型模型帮助大语言模型识别并阻止有害内容,但现有评估大多以英语为中心,忽视语言与文化多样性。现有多语言安全基准常依赖机器翻译的英文数据,无法捕捉低资源语言的细微差别。尽管东南亚地区语言多样且存在独特安全风险(如敏感政治言论、区域性虚假信息),其语言仍严重被忽视。解决这一缺口需要反映本地规范与危害场景的原生基准。我们提出SEA-SafeguardBench,首个面向东南亚的人工验证安全基准,涵盖8种语言、21,640个样本,分为通用、真实世界和内容生成三类子集。实验结果表明,即使最先进的大模型与防护机制,在东南亚文化与危害场景中也面临挑战,相较于英文文本表现明显下降。
原文摘要 · Abstract (English)
Safeguard models help large language models (LLMs) detect and block harmful content, but most evaluations remain English-centric and overlook linguistic and cultural diversity. Existing multilingual safety benchmarks often rely on machine-translated English data, which fails to capture nuances in low-resource languages. Southeast Asian (SEA) languages are underrepresented despite the region's linguistic diversity and unique safety concerns, from culturally sensitive political speech to region-specific misinformation. Addressing these gaps requires benchmarks that are natively authored to reflect local norms and harm scenarios. We introduce SEA-SafeguardBench, the first human-verified safety benchmark for SEA, covering eight languages, 21,640 samples, across three subsets: general, in-the-wild, and content generation. The experimental results from our benchmark demonstrate that even state-of-the-art LLMs and guardrails are challenged by SEA cultural and harm scenarios and underperform when compared to English texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。