arXiv:2409.10997cs.CL2024-09

测试Transformer模型在七类噪声下的问答鲁棒性。

Contextual Breach: Assessing the Robustness of Transformer-based QA Models

  • 在SQuAD数据集上下文中加入七种噪声,每种五级强度。
  • 发现Transformer模型在真实文本输入下存在明显脆弱性。
  • 适合关注模型安全与实际应用可靠性的研究者。

基于上下文的问答模型容易受到输入上下文中的对抗性扰动影响,这类扰动在现实场景中常见。本文构建了一个独特数据集,在SQuAD数据集的上下文中引入七种不同类型的对抗性噪声,每种噪声设置五个强度级别。为量化模型鲁棒性,采用标准化的鲁棒性指标评估模型在不同噪声类型和强度下的表现。对基于Transformer的问答模型进行实验,揭示了其在真实文本输入下的鲁棒性缺陷,并提供了重要洞察。

原文摘要 · Abstract (English)

Contextual question-answering models are susceptible to adversarial perturbations to input context, commonly observed in real-world scenarios. These adversarial noises are designed to degrade the performance of the model by distorting the textual input. We introduce a unique dataset that incorporates seven distinct types of adversarial noise into the context, each applied at five different intensity levels on the SQuAD dataset. To quantify the robustness, we utilize robustness metrics providing a standardized measure for assessing model performance across varying noise types and levels. Experiments on transformer-based question-answering models reveal robustness vulnerabilities and important insights into the model's performance in realistic textual input.

问答系统对抗攻击鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。