提升多语言AI安全能力,防范跨语言越狱风险
Towards Safe Multilingual Frontier AI
- 测试5个先进模型在24种欧盟官方语言下的安全表现
- 发现资源少的语言更易被跨语言越狱攻击
- 建议强制评估多语言能力,推动政策支持
语言包容性大模型需在不同语言提示下保持良好性能,以促进全球AI普惠。依赖语言翻译的多语言越狱攻击会削弱AI系统的安全性与包容性。本文通过测试五个先进AI模型在24种欧盟官方语言中的表现,分析语言资源水平与模型受多语言越狱攻击脆弱性的关系。基于已有研究,提出符合欧盟法律框架的政策建议,包括强制开展多语言能力与漏洞评估、开展公众意见调研、以及国家对多语言AI发展的支持。这些措施旨在通过欧盟政策推动提升AI安全与功能,支持《欧盟人工智能法案》实施,并为欧洲人工智能办公室监管工作提供依据。
原文摘要 · Abstract (English)
Linguistically inclusive LLMs -- which maintain good performance regardless of the language with which they are prompted -- are necessary for the diffusion of AI benefits around the world. Multilingual jailbreaks that rely on language translation to evade safety measures undermine the safe and inclusive deployment of AI systems. We provide policy recommendations to enhance the multilingual capabilities of AI while mitigating the risks of multilingual jailbreaks. We examine how a language's level of resourcing relates to how vulnerable LLMs are to multilingual jailbreaks in that language. We do this by testing five advanced AI models across 24 official languages of the EU. Building on prior research, we propose policy actions that align with the EU legal landscape and institutional framework to address multilingual jailbreaks, while promoting linguistic inclusivity. These include mandatory assessments of multilingual capabilities and vulnerabilities, public opinion research, and state support for multilingual AI development. The measures aim to improve AI safety and functionality through EU policy initiatives, guiding the implementation of the EU AI Act and informing regulatory efforts of the European AI Office.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。