arXiv:2511.11309cs.CL2025-11

首个针对孟加拉语模型的语义保持攻击与防御基准,提升模型鲁棒性。

destroR: A Benchmark and Adversarial-Training Defense for Bangla Transfer Models under Meaning-Preserving Attacks

  • 设计三类语义不变的孟加拉语攻击方法,保持句子流畅且语义一致。
  • 在四个数据集上测试五种模型,发现词替换基线攻击成功率高达54.2%。
  • 通过对抗训练显著提升所有模型鲁棒性,多语言模型反而更抗攻击。

基于Transformer的迁移模型如今主导孟加拉语情感分类任务,但其对抗鲁棒性尚未被充分研究,且此前无研究将孟加拉语攻击套件与可量化解法结合。本文提出destroR,一个统一的评估与强化孟加拉语文本分类器的流程。首先,引入三种语义保持型孟加拉语攻击方法:改写攻击、反向翻译攻击和独热词替换攻击,扰动输入的同时生成流畅且语义忠实的句子,引发模型预测困惑而非单纯噪声。其次,构建鲁棒性基准,评估五种模型(BanglaBERT、BanglishBERT、XLM-RoBERTa、MuRIL、IndicBERTv2)在四个数据集上对五种攻击的表现,并与TextFooler和BAE两种强基线在相同协议下对比。第三,通过联合所有攻击家族进行对抗训练,报告完整的鲁棒性矩阵。分析显示:词替换基线攻击比语义约束方法更强(BAE达54.2%攻击成功率);在所有攻击族上对抗训练均有效降低残余攻击成功率;出乎意料的是,多语言模型MuRIL的鲁棒性显著优于专用孟加拉语模型。所有模型、对抗数据及代码均已公开,确保完全可复现。

原文摘要 · Abstract (English)

Transformer-based transfer models now dominate Bangla sentiment classification, yet their adversarial robustness remains largely unexamined, and no prior study pairs a Bangla attack suite with a defense that measurably recovers robustness. We address this gap with destroR, a unified pipeline for evaluating and hardening Bangla text classifiers. First, we introduce three meaning-preserving Bangla attack recipes a paraphrase attack, a back-translation attack, and a one-hot word-swap attack that perturb inputs while regenerating fluent, semantically faithful sentences, inducing model prediction perplexity rather than input noise. Second, we construct a robustness benchmark that evaluates five transfer models (BanglaBERT, BanglishBERT, XLM-RoBERTa, MuRIL, and IndicBERTv2) across four datasets against five attacks, placing our recipes against two strong word-substitution baselines, TextFooler and BAE, under an identical protocol. Third, we harden every model through adversarial training and report a full robustness matrix. Our analysis yields three findings: word-substitution baselines are more potent than semantically constrained recipes (BAE reaches a 54.2% attack success rate); adversarial training on the union of all attack families lowers residual attack success for every attack; and, contrary to expectation, the Indic-multilingual MuRIL backbone is markedly more robust than the Bangla-dedicated models. All models, adversarial data, and code are released for full reproducibility.

对抗攻击孟加拉语文本分类鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。