跨语言AI安全评估实验揭示模型在不同语言中表现差异。
Improving Methodologies for LLM Evaluations Across Global Languages
- 十种语言测试开放模型,用LLM与人工双评估
- 6000+新译提示,发现安全防护能力语言间差异
- 提出文化适配翻译和清晰标注指南等改进方法
随着前沿AI模型在全球部署,其在多元语言文化背景下的安全性与可靠性至关重要。由新加坡AI安全研究所牵头,国际先进AI测评科学网络成员(涵盖新加坡、日本、澳大利亚、加拿大、欧盟、法国、肯尼亚、韩国及英国)联合开展多语言评估。测试覆盖粤语、英语、波斯语、法语、日语、韩语、斯瓦希里语、马来语、中文普通话及泰卢固语共十种语言,包含高资源与低资源语言。通过6000余条新翻译提示,在隐私、非暴力犯罪、暴力犯罪、知识产权及越狱鲁棒性五类危害上,采用大模型作为评判者与人工标注相结合的方式进行评估。结果显示,模型安全行为在不同语言和危害类型间存在显著差异,且评估者可靠性(大模型与人工)亦有变化。研究还提炼出改进多语言安全评估的方法论建议,包括文化适配的翻译策略、压力测试型评估提示及更清晰的人类标注规范。该工作为构建全球统一的高级AI系统多语言安全测试框架迈出第一步,并呼吁学术界与产业界持续合作。
原文摘要 · Abstract (English)
As frontier AI models are deployed globally, it is essential that their behaviour remains safe and reliable across diverse linguistic and cultural contexts. To examine how current model safeguards hold up in such settings, participants from the International Network for Advanced AI Measurement, Evaluation and Science, including representatives from Singapore, Japan, Australia, Canada, the EU, France, Kenya, South Korea and the UK conducted a joint multilingual evaluation exercise. Led by Singapore AISI, two open-weight models were tested across ten languages spanning high and low resourced groups: Cantonese English, Farsi, French, Japanese, Korean, Kiswahili, Malay, Mandarin Chinese and Telugu. Over 6,000 newly translated prompts were evaluated across five harm categories (privacy, non-violent crime, violent crime, intellectual property and jailbreak robustness), using both LLM-as-a-judge and human annotation. The exercise shows how safety behaviours can vary across languages. These include differences in safeguard robustness across languages and harm types and variation in evaluator reliability (LLM-as-judge vs. human review). Further, it also generated methodological insights for improving multilingual safety evaluations, such as the need for culturally contextualised translations, stress-tested evaluator prompts and clearer human annotation guidelines. This work represents an initial step toward a shared framework for multilingual safety testing of advanced AI systems and calls for continued collaboration with the wider research community and industry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。