arXiv:2409.05021cs.CL2024-09IJCAI被引 4

融合视觉信息生成更隐蔽、更强的翻译攻击文本

Vision-fused Attack: Advancing Aggressive and Stealthy Adversarial Text against Neural Machine Translation

  • 结合视觉特征扩展语义空间,提升攻击强度
  • 基于人阅读习惯筛选文本,显著增强隐蔽性
  • 适用于评估和强化大模型翻译安全性

尽管神经机器翻译(NMT)在日常应用中取得成功,但其对对抗攻击存在脆弱性。现有研究在攻击能力与人类不可察觉性方面均不足,因仅关注语言层面。本文提出视觉融合攻击(VFA)框架,通过视觉融合策略扩展语义解空间,提升攻击候选的攻击能力;同时设计感知保留的对抗文本选择策略,使最终生成文本更符合人类阅读习惯,更具欺骗性。在多种模型(包括LLaMA、GPT-3.5等大语言模型)上的实验表明,VFA相比基线有显著优势,平均攻击成功率(ASR)提升达81%,结构相似性(SSIM)提升14%。

原文摘要 · Abstract (English)

While neural machine translation (NMT) models achieve success in our daily lives, they show vulnerability to adversarial attacks. Despite being harmful, these attacks also offer benefits for interpreting and enhancing NMT models, thus drawing increased research attention. However, existing studies on adversarial attacks are insufficient in both attacking ability and human imperceptibility due to their sole focus on the scope of language. This paper proposes a novel vision-fused attack (VFA) framework to acquire powerful adversarial text, i.e., more aggressive and stealthy. Regarding the attacking ability, we design the vision-merged solution space enhancement strategy to enlarge the limited semantic solution space, which enables us to search for adversarial candidates with higher attacking ability. For human imperceptibility, we propose the perception-retained adversarial text selection strategy to align the human text-reading mechanism. Thus, the finally selected adversarial text could be more deceptive. Extensive experiments on various models, including large language models (LLMs) like LLaMA and GPT-3.5, strongly support that VFA outperforms the comparisons by large margins (up to 81%/14% improvements on ASR/SSIM).

对抗攻击机器翻译视觉融合大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。