arXiv:2509.15260cs.CL2025-09EMNLP被引 3

针对新加坡低资源语言,构建了新型安全评测基准。

Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages

  • 采用红队测试法,评估大模型在对话等三场景下的安全漏洞。
  • 发现主流多语言模型在方言和多语种环境中的防护能力严重不足。
  • 适合关注AI伦理、跨文化安全与低资源语言研究者参考。

大型语言模型的进展虽已显著推动自然语言处理,但其在低资源多语种环境中的安全性仍缺乏深入探索。本文旨在填补这一空白,提出 extsf{SGToxicGuard}——一个面向新加坡多元语言生态(包括新加坡式英语、中文、马来语、泰米尔语)的新数据集与评估框架,用于系统性地测试大模型在真实场景下的安全表现,涵盖对话、问答与内容生成三类任务。通过大规模实验,我们揭示了当前先进多语言模型在文化敏感性和毒性防御方面的关键缺陷。研究成果为构建更具包容性与安全性的多语种AI系统提供了重要依据。数据集可访问:https://github.com/Social-AI-Studio/SGToxicGuard。

原文摘要 · Abstract (English)

The advancement of Large Language Models (LLMs) has transformed natural language processing; however, their safety mechanisms remain under-explored in low-resource, multilingual settings. Here, we aim to bridge this gap. In particular, we introduce \textsf{SGToxicGuard}, a novel dataset and evaluation framework for benchmarking LLM safety in Singapore's diverse linguistic context, including Singlish, Chinese, Malay, and Tamil. SGToxicGuard adopts a red-teaming approach to systematically probe LLM vulnerabilities in three real-world scenarios: \textit{conversation}, \textit{question-answering}, and \textit{content composition}. We conduct extensive experiments with state-of-the-art multilingual LLMs, and the results uncover critical gaps in their safety guardrails. By offering actionable insights into cultural sensitivity and toxicity mitigation, we lay the foundation for safer and more inclusive AI systems in linguistically diverse environments.\footnote{Link to the dataset: https://github.com/Social-AI-Studio/SGToxicGuard.} \textcolor{red}{Disclaimer: This paper contains sensitive content that may be disturbing to some readers.}

大模型安全多语言红队测试AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。