提出可扩展的敏感度评估框架,用于分析文本分类中关键词影响。
SMAB: MAB based word Sensitivity Estimation Framework and its Applications in Adversarial Text Generation
- 基于多臂赌博机设计高效敏感度估算方法
- 在无真实标签时,敏感度可替代准确率作为评估指标
- 提升对抗文本生成成功率15.58%,优于现有方法
为理解序列分类任务的复杂性,Hahn等(2021)将敏感度定义为可独立改变以影响输出的输入子集数量。尽管有效,该方法因时间复杂度呈指数增长,难以大规模应用。为此,本文提出基于多臂赌博机的敏感度评估框架(SMAB),可高效计算任意数据集上文本分类器的词级局部(句子级)与全局(聚合)敏感度。通过多项应用验证其有效性:在CHECKLIST生成的情感分析数据集上,算法成功识别出直观高敏感与低敏感词汇;在多个任务与语言上,敏感度可作为无真实标签情况下的准确率代理指标;在对抗样本生成中,使用敏感度引导扰动提示使攻击成功率提升15.58%;在对抗改写生成中,引入敏感度作为额外奖励,相较当前最优方法提升12.00%。
原文摘要 · Abstract (English)
To understand the complexity of sequence classification tasks, Hahn et al. (2021) proposed sensitivity as the number of disjoint subsets of the input sequence that can each be individually changed to change the output. Though effective, calculating sensitivity at scale using this framework is costly because of exponential time complexity. Therefore, we introduce a Sensitivity-based Multi-Armed Bandit framework (SMAB), which provides a scalable approach for calculating word-level local (sentence-level) and global (aggregated) sensitivities concerning an underlying text classifier for any dataset. We establish the effectiveness of our approach through various applications. We perform a case study on CHECKLIST generated sentiment analysis dataset where we show that our algorithm indeed captures intuitively high and low-sensitive words. Through experiments on multiple tasks and languages, we show that sensitivity can serve as a proxy for accuracy in the absence of gold data. Lastly, we show that guiding perturbation prompts using sensitivity values in adversarial example generation improves attack success rate by 15.58%, whereas using sensitivity as an additional reward in adversarial paraphrase generation gives a 12.00% improvement over SOTA approaches. Warning: Contains potentially offensive content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。