用少量字符改动和小模型,就能攻破低资源语言的LLM安全机制。
Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models
- 仅改几个字符,结合小代理模型计算词重要性,低成本生成攻击。
- 在波兰语中成功使多个LLM预测结果大幅偏离,暴露安全漏洞。
- 方法可推广至其他低资源语言,适合安全评测与模型鲁棒性研究者。
近年来大型语言模型(LLMs)在多种自然语言处理任务中表现出色,但其对越狱攻击和扰动的敏感性仍需额外评估。尽管许多LLM具备多语言能力,但安全训练数据主要集中在英语等高资源语言,导致其在波兰语等低资源语言上可能面临安全漏洞。本文展示,仅通过修改少数字符,并利用小型代理模型计算词重要性,即可低成本生成强大攻击。实验表明,这类字符级与词级攻击能显著改变不同LLM的输出预测,暴露出其内部安全机制的潜在缺陷。我们在波兰语上验证了该攻击方法的有效性,发现其在低资源语言中的普遍脆弱性。此外,该方法可扩展至其他语言。论文已公开相关数据集与代码,以促进后续研究。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated impressive capabilities across various natural language processing (NLP) tasks in recent years. However, their susceptibility to jailbreaks and perturbations necessitates additional evaluations. Many LLMs are multilingual, but safety-related training data contains mainly high-resource languages like English. This can leave them vulnerable to perturbations in low-resource languages such as Polish. We show how surprisingly strong attacks can be cheaply created by altering just a few characters and using a small proxy model for word importance calculation. We find that these character and word-level attacks drastically alter the predictions of different LLMs, suggesting a potential vulnerability that can be used to circumvent their internal safety mechanisms. We validate our attack construction methodology on Polish, a low-resource language, and find potential vulnerabilities of LLMs in this language. Additionally, we show how it can be extended to other languages. We release the created datasets and code for further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。