arXiv:2507.03473cs.CLcs.AI2025-07被引 3

关注低资源语言的NLP安全,发现模型规模与多语能力不保证安全

Beyond Weaponization: NLP Security for Medium and Lower-Resourced Languages in Their Own Right

  • 扩展70种语言的对抗攻击,评估中低资源语言模型安全
  • 单语模型参数量不足,安全难以保障;多语化也不总能提升安全
  • 为低资源语言社区提供更可靠的模型部署思路

尽管已有大量证据表明多语言特性可能被用于攻击语言模型(LMs),但当前NLP安全研究仍以英语为主。在模型安全领域,'英语优先'的惯例与网络安全中需预判最坏情况的标准相悖。为应对最坏情况,研究者应关注语言模型安全中的薄弱环节——中低资源语言。本文针对中低资源语言的模型安全展开研究,将现有对抗攻击方法扩展至最多70种语言,评估单语和多语语言模型的安全性。分析发现,单语模型因参数总量过少,难以保障安全;而多语化虽有帮助,却无法确保安全性的提升。这些结果强调了在部署语言模型时,必须重视中低资源语言社区的安全需求。

原文摘要 · Abstract (English)

Despite mounting evidence that multilinguality can be easily weaponized against language models (LMs), works across NLP Security remain overwhelmingly English-centric. In terms of securing LMs, the NLP norm of "English first" collides with standard procedure in cybersecurity, whereby practitioners are expected to anticipate and prepare for worst-case outcomes. To mitigate worst-case outcomes in NLP Security, researchers must be willing to engage with the weakest links in LM security: lower-resourced languages. Accordingly, this work examines the security of LMs for lower- and medium-resourced languages. We extend existing adversarial attacks for up to 70 languages to evaluate the security of monolingual and multilingual LMs for these languages. Through our analysis, we find that monolingual models are often too small in total number of parameters to ensure sound security, and that while multilinguality is helpful, it does not always guarantee improved security either. Ultimately, these findings highlight important considerations for more secure deployment of LMs, for communities of lower-resourced languages.

NLP安全低资源语言对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。