arXiv:2507.05980cs.CLcs.LG2025-07被引 2

构建新加坡多语种安全评测集,揭示模型在本地方言中的安全漏洞

Lost in Localization: Building RabakBench with Human-in-the-Loop Validation to Measure Multilingual Safety Gaps

  • 用LLM生成真实有害内容,结合人工校验构建多语言安全数据集
  • 覆盖5000+样本,六类细粒度安全标注,人工一致性达0.70-0.80
  • 针对本地化方言的评测发现主流安全防护系统性能严重下降

大型语言模型在低资源语言变体(如混合语方言和地方口音)中常无法保持安全。本文提出RabakBench,一个面向新加坡多元语言环境的多语种安全基准,涵盖Singlish、中文、马来语和泰米尔语。通过三阶段流程构建:(1) 生成:利用LLM驱动的红队测试扩充真实网络有害内容;(2) 标注:采用多数投票机制的半自动化多标签标注;(3) 翻译:实现高保真、毒性保留的跨语言转换。最终数据集包含超过5,000个样本,覆盖六类细粒度安全类别。尽管使用LLM提升可扩展性,框架仍保持严格的人工监督,达成0.70-0.80的标注者间一致性。对13个前沿安全防护系统的评估显示显著性能下降,凸显本地化评测的必要性。RabakBench为服务弱势社区的安全基准建设提供可复现框架。

原文摘要 · Abstract (English)

Large language models (LLMs) often fail to maintain safety in low-resource language varieties, such as code-mixed vernaculars and regional dialects. We introduce RabakBench, a multilingual safety benchmark and scalable pipeline localized to Singapore's unique linguistic landscape, covering Singlish, Chinese, Malay, and Tamil. We construct the benchmark through a three-stage pipeline: (1) Generate: augmenting real-world unsafe web content via LLM-driven red teaming; (2) Label: applying semi-automated multi-label annotation using majority-voted LLM labelers; and (3) Translate: performing high-fidelity, toxicity-preserving translation. The resulting dataset contains over 5,000 examples across six fine-grained safety categories. Despite using LLMs for scalability, our framework maintains rigorous human oversight, achieving 0.70-0.80 inter-annotator agreement. Evaluations of 13 state-of-the-art guardrails reveal significant performance degradation, underscoring the need for localized evaluation. RabakBench provides a reproducible framework for building safety benchmarks in underserved communities.

多语种安全评测本地化语言多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。