arXiv:2503.13081cs.CLcs.AI2025-03被引 1

提出自动评估大模型多语言安全漏洞的框架,发现低资源语言易出问题但风险较低。

A Framework to Assess Multilingual Vulnerabilities of LLMs

  • 构建自动化框架,跨8种语言评估6个主流大模型的安全性
  • 低资源语言模型响应常不连贯,漏洞源于性能差而非刻意攻击
  • 结果与人工评估高度一致,适合安全研究者和模型开发者参考

大型语言模型(LLMs)正扩展多语言理解与生成能力。尽管经过安全训练以避免回答非法问题,但训练数据与人工评估资源分布不均,使模型在低资源语言(LRL)中更易受攻击。本文提出一种自动评估多语言漏洞的框架,对六种常用LLMs在八种不同资源水平的语言上进行评估。通过在两种语言中开展人工验证,确认该框架结果与人类判断基本一致。研究发现,低资源语言存在漏洞,但多数因模型性能不佳导致回应不连贯,实际风险较小。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are acquiring a wider range of capabilities, including understanding and responding in multiple languages. While they undergo safety training to prevent them from answering illegal questions, imbalances in training data and human evaluation resources can make these models more susceptible to attacks in low-resource languages (LRL). This paper proposes a framework to automatically assess the multilingual vulnerabilities of commonly used LLMs. Using our framework, we evaluated six LLMs across eight languages representing varying levels of resource availability. We validated the assessments generated by our automated framework through human evaluation in two languages, demonstrating that the framework's results align with human judgments in most cases. Our findings reveal vulnerabilities in LRL; however, these may pose minimal risk as they often stem from the model's poor performance, resulting in incoherent responses.

大模型安全多语言漏洞评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。