提出新方法生成字符级对抗样本,提升机器翻译鲁棒性
Extend Adversarial Policy Against Neural Machine Translation via Unknown Token
- 基于强化学习引入字符扰动替代词元替换
- 在基线方法失效场景下仍能生成高效对抗样本
- 适合研究模型安全与防御机制的开发者
生成对抗样本有助于提升主流神经机器翻译(NMT)的鲁棒性。然而,现有对抗策略多针对固定分词,难以应对涉及多样分词方式的字符级扰动。本文基于强化学习的对抗生成框架,提出「DexChar策略」,引入字符级扰动以替代传统的词元替换。同时,改进自监督匹配机制,为强化学习提供更符合语义约束的反馈信号。实验表明,该方法在基线对抗攻击失效的场景下依然有效,可生成高质量对抗样本,用于系统分析与优化。
原文摘要 · Abstract (English)
Generating adversarial examples contributes to mainstream neural machine translation~(NMT) robustness. However, popular adversarial policies are apt for fixed tokenization, hindering its efficacy for common character perturbations involving versatile tokenization. Based on existing adversarial generation via reinforcement learning~(RL), we propose the `DexChar policy' that introduces character perturbations for the existing mainstream adversarial policy based on token substitution. Furthermore, we improve the self-supervised matching that provides feedback in RL to cater to the semantic constraints required during training adversaries. Experiments show that our method is compatible with the scenario where baseline adversaries fail, and can generate high-efficiency adversarial examples for analysis and optimization of the system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。