探究词形变化语言在对抗攻击下的鲁棒性,揭示语法结构如何影响模型脆弱性。
Unpacking Robustness in Inflectional Languages: Adversarial Evaluation and Mechanistic Insights
- 基于边缘归因修补法设计新评估协议,分析词形变化对模型的影响。
- 在波兰语和英语的并行语料上验证攻击效果,发现词形复杂度提升模型脆弱性。
- 提出可解释的机制框架,适合关注模型安全与语言结构的研究者。
现有对抗样本生成技术多针对非屈折语言(如英语)设计,包括TextBugger等引入微小、不易察觉的词汇扰动,或TextFooler通过同义词替换保持语义但改变预测结果。本文首次系统评估对抗攻击在屈折语言中的表现。基于机制可解释性思想,提出一种受边属性修补(EAP)启发的新评估协议。利用包含屈折与合成变体的双语并行语料库(波兰语与英语),在任务导向数据集MultiEmo基础上构建新基准,识别模型内部与词形相关的机制组件,并分析其在攻击下的行为。实验揭示词形变化显著影响模型鲁棒性,为理解语言结构与模型安全性关系提供新视角。
原文摘要 · Abstract (English)
Various techniques are used in the generation of adversarial examples, including methods such as TextBugger which introduce minor, hardly visible perturbations to words leading to changes in model behaviour. Another class of techniques involves substituting words with their synonyms in a way that preserves the text's meaning but alters its predicted class, with TextFooler being a prominent example of such attacks. Most adversarial example generation methods are developed and evaluated primarily on non-inflectional languages, typically English. In this work, we evaluate and explain how adversarial attacks perform in inflectional languages. To explain the impact of inflection on model behaviour and its robustness under attack, we designed a novel protocol inspired by mechanistic interpretability, based on Edge Attribution Patching (EAP) method. The proposed evaluation protocol relies on parallel task-specific corpora that include both inflected and syncretic variants of texts in two languages -- Polish and English. To analyse the models and explain the relationship between inflection and adversarial robustness, we create a new benchmark based on task-oriented dataset MultiEmo, enabling the identification of mechanistic inflection-related elements of circuits within the model and analyse their behaviour under attack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。