arXiv:2502.16901cs.CLcs.AI2025-02被引 1

跨语言后门攻击可让单语数据污染影响多语言模型安全

Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs

  • 利用共享嵌入空间,单语言中毒数据触发跨语言后门
  • 罕见与高频词作触发器,使模型在多语言下产生隐蔽恶意响应
  • 适用于研究多语言模型安全的学者与防御开发者

我们研究多语言大模型中的跨语言后门攻击(X-BAT),揭示了在一种语言中植入的后门可通过共享嵌入空间自动迁移至其他语言。以毒性分类为例,攻击者仅需污染单一语言数据,利用罕见和高出现词作为特定有效触发器,即可破坏多语言系统。研究发现,该漏洞源于模型架构对信息流的影响,导致后门在推理过程中隐性激活。相关代码与数据已公开于 https://github.com/himanshubeniwal/X-BAT。

原文摘要 · Abstract (English)

We explore \textbf{C}ross-lingual \textbf{B}ackdoor \textbf{AT}tacks (X-BAT) in multilingual Large Language Models (mLLMs), revealing how backdoors inserted in one language can automatically transfer to others through shared embedding spaces. Using toxicity classification as a case study, we demonstrate that attackers can compromise multilingual systems by poisoning data in a single language, with rare and high-occurring tokens serving as specific, effective triggers. Our findings expose a critical vulnerability that influences the model's architecture, resulting in a concealed backdoor effect during the information flow. Our code and data are publicly available https://github.com/himanshubeniwal/X-BAT.

后门攻击多语言模型安全漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。