构建12语言安全评估基准,揭示大模型跨语言安全表现差异。
LinguaSafe: A Comprehensive Multilingual Safety Benchmark for Large Language Models
- 融合翻译、再创作与原创数据,构建45k条多语言真实语料
- 跨语言安全表现差异显著,同资源水平语言间也存在巨大鸿沟
- 提供细粒度评估框架,适合关注多语言安全对齐的研究者
大型语言模型在全球技术中的广泛应用,要求其在多种语言和文化背景下具备安全性。现有跨语言安全评估因数据不全面、覆盖不足而效果受限,难以支撑鲁棒的多语言安全对齐。为此,我们提出LinguaSafe,一个精心设计的综合性多语言安全基准。该数据集包含45,000条样本,覆盖匈牙利语至马来语等12种语言,通过翻译、再创作与原生数据相结合的方式构建,填补了从匈牙利语到马来语等低资源语言在安全评估上的空白。LinguaSafe提供多维度、细粒度的评估框架,涵盖直接与间接安全测试,以及过度敏感性评估。实验显示,不同领域与语言间的安全性和助人性表现差异显著,即使在资源水平相近的语言中亦如此。该基准提供一套完整的评估指标,强调全面评估多语言安全对齐的重要性。数据集与代码已公开,以促进多语言大模型安全研究。
原文摘要 · Abstract (English)
The widespread adoption and increasing prominence of large language models (LLMs) in global technologies necessitate a rigorous focus on ensuring their safety across a diverse range of linguistic and cultural contexts. The lack of a comprehensive evaluation and diverse data in existing multilingual safety evaluations for LLMs limits their effectiveness, hindering the development of robust multilingual safety alignment. To address this critical gap, we introduce LinguaSafe, a comprehensive multilingual safety benchmark crafted with meticulous attention to linguistic authenticity. The LinguaSafe dataset comprises 45k entries in 12 languages, ranging from Hungarian to Malay. Curated using a combination of translated, transcreated, and natively-sourced data, our dataset addresses the critical need for multilingual safety evaluations of LLMs, filling the void in the safety evaluation of LLMs across diverse under-represented languages from Hungarian to Malay. LinguaSafe presents a multidimensional and fine-grained evaluation framework, with direct and indirect safety assessments, including further evaluations for oversensitivity. The results of safety and helpfulness evaluations vary significantly across different domains and different languages, even in languages with similar resource levels. Our benchmark provides a comprehensive suite of metrics for in-depth safety evaluation, underscoring the critical importance of thoroughly assessing multilingual safety in LLMs to achieve more balanced safety alignment. Our dataset and code are released to the public to facilitate further research in the field of multilingual LLM safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。