主流大模型对非主流语言安全性能不足,论文揭示其背后成因与治理路径。
The Multilingual Divide and Its Impact on Global AI Safety
- 分析主流语言外模型能力与安全性的系统性差距
- 指出多语言数据匮乏是核心障碍之一
- 建议政策与治理层面推动多语种数据共建
近年来大语言模型能力虽有显著提升,但在全球少数主导语言之外,仍存在显著的能力与安全性差距。本文为研究者、政策制定者及治理专家提供关于弥合人工智能‘语言鸿沟’的关键挑战的概览,并探讨如何降低跨语言的安全风险。文章分析了语言鸿沟的成因及其加剧机制,揭示其如何导致全球人工智能安全水平的不平等。识别出解决这些挑战的主要障碍,并提出政策与治理实践者可通过支持多语言数据集建设、提升透明度和推动相关研究,来缓解与语言鸿沟相关的安全问题。
原文摘要 · Abstract (English)
Despite advances in large language model capabilities in recent years, a large gap remains in their capabilities and safety performance for many languages beyond a relatively small handful of globally dominant languages. This paper provides researchers, policymakers and governance experts with an overview of key challenges to bridging the "language gap" in AI and minimizing safety risks across languages. We provide an analysis of why the language gap in AI exists and grows, and how it creates disparities in global AI safety. We identify barriers to address these challenges, and recommend how those working in policy and governance can help address safety concerns associated with the language gap by supporting multilingual dataset creation, transparency, and research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。