arXiv:2605.00689cs.CLcs.CR2026-05被引 1

基于地区法规构建多语言安全评测与防护系统,提升大模型跨文化合规能力。

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models

论文配图:ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
图 1 · 摘自论文原文
  • 直接从各地法律文本提取风险类别和细则,生成文化合规的多语言数据
  • 推出1.5B与7B两个版本的防护模型,在14种语言上均超越11个基线
  • 适合需遵守本地法规的跨国AI应用开发者使用

随着大语言模型在跨语言场景中日益普及,确保其在不同监管与文化环境下的安全性成为关键挑战。现有多语言评测基准大多依赖通用风险分类和机器翻译,使防护模型局限于预设类别,难以适配地区性法规与文化差异。为此,我们提出ML-Bench,一个覆盖14种语言、基于政策原文的多语言安全评测基准。该基准直接从区域法规中提取风险类别与细粒度规则,用于生成符合当地法律与文化语境的安全数据,实现跨语言的合规评估。在此基础上,我们开发了基于扩散大语言模型(dLLM)的ML-Guard防护系统,包含1.5B轻量版用于快速安全判断,以及7B增强版支持定制化合规分析并提供详细解释。我们在6个现有基准及ML-Bench上对11个强基线进行实验,结果表明ML-Guard持续领先。我们期望ML-Bench与ML-Guard能推动面向法规与文化的多语言防护系统发展。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) are increasingly deployed in cross-linguistic contexts, ensuring safety in diverse regulatory and cultural environments has become a critical challenge. However, existing multilingual benchmarks largely rely on general risk taxonomies and machine translation, which confines guardrail models to these predefined categories and hinders their ability to align with region-specific regulations and cultural nuances. To bridge these gaps, we introduce ML-Bench, a policy-grounded multilingual safety benchmark covering 14 languages. ML-Bench is constructed directly from regional regulations, where risk categories and fine-grained rules derived from jurisdiction-specific legal texts are directly used to guide the generation of multilingual safety data, enabling culturally and legally aligned evaluation across languages. Building on ML-Bench, we develop ML-Guard, a Diffusion Large Language Model (dLLM)-based guardrail model that supports multilingual safety judgment and policy-conditioned compliance assessment. ML-Guard has two variants, one 1.5B lightweight model for fast `safe/unsafe' checking and a more capable 7B model for customized compliance checking with detailed explanations. We conduct extensive experiments against 11 strong guardrail baselines across 6 existing multilingual safety benchmarks and our ML-Bench, and show that ML-Guard consistently outperforms prior methods. We hope that ML-Bench and ML-Guard can help advance the development of regulation-aware and culturally aligned multilingual guardrail systems.

大模型安全多语言合规防护政策对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。